# Artifact Registries: Docker Hub, Nexus, Artifactory

> Where your builds live: container and package registries that store, version, and serve your artifacts - with the proxying and access control teams rely on.


---

# Artifact Registries: Docker Hub, Nexus, Artifactory

Your build finishes, your tests go green, and then you hit the quiet question nobody documented: where does the *output* go? The image, the jar, the npm package - they need a home that the next pipeline stage, the next teammate, and production can all pull from. That home is an artifact registry, and the difference between treating it as a dumb bucket and understanding how it really works is the difference between reliable deploys and the day a moved `latest` tag silently ships old code.

This guide gives you the mental model first - what an artifact is, why registries exist, and the one rule (immutability) that everything else hangs on. Then the everyday flow of pushing and pulling across Docker Hub, GHCR, Nexus, and Artifactory. Then the production reality: proxying public registries for speed and resilience, private packages, retention, scanning, and exactly why the public-tag-overwrite trap bites teams who thought a tag was a promise.

## How to read this

- **Want it to actually click?** Read in order. Phase 1 installs the immutability mental model that phases 2 and 3 lean on the whole way through.
- **Already pushing images and only need the gotchas?** Phase 3 is the production phase - proxy repos, retention, scanning, and the tag-overwrite trap. But skim phase 1's immutability section first; it is the reason the trap exists.

## The phases

1. **[What a Registry Actually Is](01-what-a-registry-actually-is.md)** - artifacts need a home; what a registry stores, how a tag points at an immutable digest, and why that one distinction governs everything else.
2. **[Pushing, Pulling, and Private Packages](02-pushing-pulling-private-packages.md)** - the everyday flow across Docker Hub, GHCR, Nexus, and Artifactory: log in, tag, push, pull, and serve private npm/Maven/PyPI packages from one place.
3. **[Proxying, Retention, Scanning, and the Tag Trap](03-proxying-retention-scanning-gotchas.md)** - proxying public registries for speed and resilience, retention so you don't drown in old builds, vulnerability scanning, and why the mutable public tag bites.


---

# What a Registry Actually Is

Think about what a build produces. You run a `docker build` and get an image. You run `mvn package` and get a jar. You run `npm publish` and you've got a tarball. These are **artifacts** - the compiled, packaged, ready-to-ship outputs of your work. They are not source code; source lives in Git. Artifacts are what source *becomes*, and they need a different kind of home, because they are large, binary, versioned, and pulled by machines you'll never log into.

That home is an **artifact registry**: a server that stores artifacts, organizes them by name and version, and serves them over HTTP to anyone (or any pipeline) with permission to pull. Git is for the recipe; the registry is for the cake.

## Two flavors, one idea

You'll meet registries in two shapes, and it helps to see they're the same idea wearing different clothes:

- **Container registries** store OCI images. Docker Hub is the public default; **GHCR** (GitHub Container Registry, `ghcr.io`) lives next to your repos; every cloud has one (ECR, GCR/Artifact Registry, ACR).
- **Package/multi-format registries** store language packages - npm, Maven, PyPI, NuGet, Go modules, and often container images too. **Nexus** (Sonatype) and **Artifactory** (JFrog) are the two big self-hosted, multi-format players. One server, many package types.

```text
  your build  ─push─▶   REGISTRY   ─pull─▶  CI / teammates / production
   (image,              (stores +
    jar, npm)            versions +
                         serves)
```

*What just happened:* the registry sits in the middle of everything that produces and everything that consumes artifacts. That central position is exactly why getting it right matters - a flaky or untrustworthy registry breaks builds and deploys everywhere downstream at once.

## The part that governs everything: tags vs digests

Here is the single most important idea in the whole guide, and the source of most registry pain when it's misunderstood.

When you push a container image, the registry stores the actual bytes under a **digest** - a content hash that looks like `sha256:9b2a...`. The digest *is* the content. Change one byte of the image and you get a completely different digest. A digest is permanent and immutable: `sha256:9b2a...` will forever mean exactly those bytes, or it won't exist at all.

A **tag** - like `1.4.0` or `latest` - is a separate, human-friendly *label* that points at a digest. And a tag is a pointer that can be moved.

```console
$ docker pull nginx:1.27.0
1.27.0: Pulling from library/nginx
Digest: sha256:67682bda769fae1ccf5183192b8daf37b64cae99c6c3302650f6f8bf5f0f95df
Status: Downloaded newer image for nginx:1.27.0
```

*What just happened:* you asked for the tag `1.27.0`, and the registry resolved it to a digest and handed you those exact bytes. The tag was the address; the digest was the content. To pin truly immutable, you can pull by digest directly - `docker pull nginx@sha256:67682b...` - and you will get those bytes no matter what anyone does to the tag later.

> The mental model that saves you: a **digest is a fact** (this content, forever). A **tag is an opinion** (right now, I think `1.27.0` means this). Opinions can change. Facts cannot.

## Why immutability is the whole game

Once you internalize "tag is a movable pointer, digest is fixed content," the rest of registry behavior stops surprising you.

Most version tags *should* behave as if immutable: once `myapp:1.4.0` is built and pushed, it should mean that exact build forever, so that the image your CI tested is byte-for-byte the image production runs. If someone can re-push a different image to `1.4.0`, you've lost the thread that connects "what we tested" to "what we shipped."

The famous exception is `latest`. It is a tag like any other, with zero special meaning to the registry - it's a conventional name people move to point at "the most recent build." That convenience is also a footgun, which we'll dissect in phase 3.

```console
# Two ways to reference the same image:
$ docker pull myapp:1.4.0                    # by tag - convenient, movable
$ docker pull myapp@sha256:9b2a3c...         # by digest - pinned, immutable
```

*What just happened:* the first line trusts that the `1.4.0` pointer still aims where you expect. The second line bypasses tags entirely and demands specific content. Production deploys that care about reproducibility increasingly pin by digest for exactly this reason.

## Why not keep artifacts in Git or a file share?

People try. Here's why it falls apart, so you understand what a registry is actually buying you:

- **Versioning that machines understand.** A registry exposes a standard API (the OCI distribution spec for images, ecosystem-specific protocols for npm/Maven/PyPI) so `docker pull`, `npm install`, and `mvn` all work against it without custom glue.
- **Deduplication by content.** Images are built in layers; packages share dependencies. A registry stores each unique layer/blob once and references it from many artifacts, so ten images sharing a base layer don't cost ten copies.
- **Access control and provenance.** Who can push? Who can pull? When was this published, and by whom? A registry answers these; a shared folder does not.

A Git repo is built for diffable text and small files with full history. A 900 MB image with binary layers is everything Git is bad at.

## For builders

If you're standing up infrastructure, your first registry decision is usually "public-hosted vs self-hosted." Docker Hub and GHCR are managed and zero-ops for images. Nexus and Artifactory are servers you run, earning their keep when you need one place for *many* package formats, fine-grained access control, and the proxying and retention features we'll cover later. There's no universally right answer - there's the one that fits your team's formats, scale, and tolerance for running infrastructure.

```quiz
[
  {
    "q": "In a container registry, what is the relationship between a tag and a digest?",
    "choices": [
      "A tag and a digest are two names for the exact same immutable thing",
      "A digest is a content hash that permanently identifies specific bytes; a tag is a movable label that points at a digest",
      "A tag is the content hash; a digest is the human-friendly label",
      "Digests are only used by Nexus, not Docker Hub"
    ],
    "answer": 1,
    "explain": "The digest is the content (a permanent hash of the bytes). The tag is a separate pointer that can be moved to point at a different digest."
  },
  {
    "q": "Why store images and packages in a registry instead of a Git repo or shared folder?",
    "choices": [
      "Registries are always cheaper than Git",
      "Git cannot store any binary files at all",
      "Registries speak standard pull APIs, deduplicate by content (shared layers/blobs), and enforce access control and provenance",
      "Registries automatically rewrite your source code"
    ],
    "answer": 2,
    "explain": "Registries are purpose-built for versioned binary artifacts: a standard pull API, content-addressed dedup, and access control - all things Git and file shares handle poorly."
  },
  {
    "q": "What does pulling by digest (image@sha256:...) guarantee that pulling by tag does not?",
    "choices": [
      "A faster download",
      "You get exactly those bytes regardless of whether anyone later moves the tag",
      "The image is automatically scanned for vulnerabilities",
      "The registry deletes all other tags"
    ],
    "answer": 1,
    "explain": "A digest is content-addressed and immutable, so it always resolves to the same bytes. A tag is a movable pointer that someone could repoint."
  }
]
```


---

# Pushing, Pulling, and Private Packages

You have the mental model. Now the muscle memory: the four moves you'll repeat thousands of times - **log in, tag, push, pull** - and how the same rhythm carries from container images to private npm and Maven packages. The commands differ per tool, but the shape is identical, and once you see the shape you stop memorizing and start understanding.

## The full name of an image tells you where it lives

Before you push anything, understand how a registry decides *where* an image goes. The full reference is structured:

```text
ghcr.io/acme/checkout-api:1.4.0
└──┬──┘ └─┬─┘ └────┬─────┘ └─┬─┘
registry  namespace  name    tag
 host    (org/user) (repo)
```

*What just happened:* the leading hostname is what routes the push. No hostname means Docker Hub by default - `nginx` is really `docker.io/library/nginx`, and `acme/checkout` is `docker.io/acme/checkout`. The moment you put `ghcr.io/...` or `nexus.internal:8082/...` in front, you're aiming at a different registry. Most "why did this push to the wrong place" confusion is a missing or wrong hostname.

## Container images: log in, tag, push, pull

Here's the complete loop against GHCR. The pattern is the same for Docker Hub, ECR, or a self-hosted Nexus - only the hostname and how you authenticate change.

```console
# 1. Authenticate (a Personal Access Token piped to stdin, never on the CLI)
$ echo "$GHCR_TOKEN" | docker login ghcr.io -u my-username --password-stdin
Login Succeeded

# 2. Tag your locally-built image with the full destination reference
$ docker tag checkout-api:dev ghcr.io/acme/checkout-api:1.4.0

# 3. Push it
$ docker push ghcr.io/acme/checkout-api:1.4.0
The push refers to repository [ghcr.io/acme/checkout-api]
5f70bf18a086: Pushed
1.4.0: digest: sha256:9b2a3c... size: 1779

# 4. Anywhere with pull access:
$ docker pull ghcr.io/acme/checkout-api:1.4.0
```

*What just happened:* `docker login` cached a credential for that host. `docker tag` gave your local image a name pointing at GHCR (tagging is local and free - it merely adds a label). `docker push` uploaded the layers and registered the `1.4.0` tag against the resulting digest. Note the `--password-stdin` form: passing a token as a CLI argument leaks it into your shell history and the process list, so always pipe it in.

> Push only the layers the registry doesn't already have. If a base layer is already there from an earlier push, you'll see `Layer already exists` instead of an upload - that's the content-addressed dedup from phase 1 saving you bandwidth.

## Tag the same image more than once

A single image (one digest) can wear several tags at once. This is how teams point both a precise version and a moving label at the same build:

```console
$ docker tag checkout-api:dev ghcr.io/acme/checkout-api:1.4.0
$ docker tag checkout-api:dev ghcr.io/acme/checkout-api:1.4
$ docker tag checkout-api:dev ghcr.io/acme/checkout-api:latest
$ docker push --all-tags ghcr.io/acme/checkout-api
```

*What just happened:* one set of bytes now answers to `1.4.0`, `1.4`, and `latest`. Pulling any of the three gives the identical digest *today*. The danger - which phase 3 unpacks - is that `1.4` and `latest` are designed to move to a *newer* build later, while `1.4.0` is the one you promise never moves.

## Private packages: one registry for npm, Maven, PyPI

This is where Nexus and Artifactory earn their place. Instead of publishing internal libraries to the public npm or Maven Central, you publish them to *your* registry, and your tooling pulls from there. The publishing protocol is each ecosystem's native one - you don't learn a new tool, you point the tool you already use at a new URL.

**npm** - point the client at your registry and authenticate via `.npmrc`:

```text
# .npmrc - scope @acme to the private registry, leave everything else public
@acme:registry=https://nexus.internal/repository/npm-private/
//nexus.internal/repository/npm-private/:_authToken=${NPM_TOKEN}
```

```console
$ npm publish               # publishes @acme/* to nexus-private
$ npm install @acme/ui-kit  # resolves @acme/* from nexus, the rest from public npm
```

*What just happened:* the scoped line says "anything named `@acme/...` comes from our private registry"; unscoped packages still flow from the public default. One registry serves your private code without you forking the entire npm ecosystem. The `_authToken` references an env var so the secret stays out of the committed file.

**Maven** - declare the repository and credentials, then deploy:

```xml
<!-- pom.xml -->
<distributionManagement>
  <repository>
    <id>acme-releases</id>
    <url>https://nexus.internal/repository/maven-releases/</url>
  </repository>
</distributionManagement>
```

```console
$ mvn deploy    # uploads the jar + pom to maven-releases
```

*What just happened:* `mvn deploy` pushed your built jar and its metadata to the Nexus `maven-releases` repo. Credentials live in `~/.m2/settings.xml` keyed by the `<id>`, so the secret stays out of the project's `pom.xml`. A teammate who lists your Nexus as a repository now resolves your internal jar exactly like a public one.

## Authentication, in one breath

Every registry, regardless of format, gates push and pull behind credentials. The mechanism varies but the principle is constant: **machines authenticate with tokens, not passwords.**

```console
# Docker Hub / GHCR / Nexus / Artifactory all follow this shape:
$ echo "$TOKEN" | docker login <registry-host> -u <user> --password-stdin

# Cloud registries often mint a short-lived token from your existing identity:
$ aws ecr get-login-password | docker login --password-stdin <acct>.dkr.ecr.<region>.amazonaws.com
```

*What just happened:* the first form is a static token you create in the registry's UI. The second is better where available - the cloud CLI exchanges your already-authenticated identity for a short-lived token, so there's no long-lived secret sitting in a CI variable waiting to leak. Prefer short-lived, scoped tokens over personal passwords everywhere you can.

## In the wild

A typical mid-size team runs one Nexus or Artifactory as the single front door: private npm under `@company`, internal Maven jars, Python wheels, *and* container images, all behind the same SSO and token policy. Developers point npm/Maven/pip/Docker at it once in their config and forget it exists - which is exactly the goal. The registry becomes invisible plumbing, and invisible plumbing is plumbing that works.

```quiz
[
  {
    "q": "In the reference ghcr.io/acme/checkout-api:1.4.0, which part decides which registry server the push goes to?",
    "choices": [
      "checkout-api (the repo name)",
      "ghcr.io (the leading hostname)",
      "1.4.0 (the tag)",
      "acme (the namespace)"
    ],
    "answer": 1,
    "explain": "The leading hostname routes the push. No hostname defaults to Docker Hub (docker.io); ghcr.io or a Nexus host aims elsewhere."
  },
  {
    "q": "Why pipe a token via --password-stdin instead of passing it as a CLI argument to docker login?",
    "choices": [
      "It makes the login faster",
      "CLI arguments leak the token into shell history and the process list; stdin keeps it out",
      "Docker only accepts tokens via stdin",
      "It compresses the token"
    ],
    "answer": 1,
    "explain": "A token passed as an argument shows up in shell history and process listings. Piping it to stdin avoids that exposure."
  },
  {
    "q": "What does the .npmrc line `@acme:registry=https://nexus.internal/...` accomplish?",
    "choices": [
      "It moves the entire public npm registry to Nexus",
      "It routes only @acme-scoped packages to the private registry, leaving unscoped packages on public npm",
      "It deletes public packages from your cache",
      "It disables authentication for @acme packages"
    ],
    "answer": 1,
    "explain": "Scoped config routes @acme/* to the private registry while everything else still resolves from the public default - one registry for your private code, no forking npm."
  }
]
```


---

# Proxying, Retention, Scanning, and the Tag Trap

Everything works in development. Then production reality arrives: a public registry rate-limits your CI mid-deploy, your storage bill quietly triples from years of dead builds, a security audit asks "are any of these images vulnerable?", and one afternoon a deploy ships old code because a tag moved under you. This phase is the bill that comes due for running registries at scale - and the well-worn answers to each item.

## Proxy public registries: speed and a seatbelt

Pulling `node:20` straight from Docker Hub on every CI run is slow and fragile: you pay the download latency every time, and if Docker Hub is down or rate-limiting, your pipeline is down with it. Public registries enforce pull rate limits, and a busy CI farm behind one shared IP hits them fast.

The fix is a **proxy (remote) repository** in Nexus or Artifactory. It sits between you and the public registry, caching what you pull:

```text
   first pull:   CI ──▶ Nexus proxy ──▶ Docker Hub   (cache miss, fetches + stores)
   later pulls:  CI ──▶ Nexus proxy ──X              (cache hit, served locally)
```

```console
# Point Docker at the proxy instead of Docker Hub directly:
$ docker pull nexus.internal:8082/docker-hub-proxy/library/node:20
```

*What just happened:* the first pull missed the cache, so Nexus fetched `node:20` from Docker Hub, stored it, and handed it to you. Every later pull is served from Nexus at LAN speed and never touches Docker Hub - so you stop burning rate-limit budget and your builds keep working even when the upstream is having a bad day. The same trick works for npm, PyPI, and Maven Central: one proxy per ecosystem, and your whole org pulls through it.

> A proxy repo is also a quiet supply-chain control point. Because everything external flows through one door, you can scan it, audit it, and (with the right tooling) block known-bad packages before they reach a build. See /guides/supply-chain-security for where this fits the bigger picture.

## Group repositories: one URL to rule them

Telling every developer "use the proxy for public, the private repo for internal" is fragile. **Group (virtual) repositories** merge several repos behind a single URL:

```text
  npm-group  =  [ npm-private  +  npm-proxy(public) ]
                        │              │
                  your @acme code   the public npm world
```

*What just happened:* developers configure exactly one registry URL - the group - and the registry resolves each request to the right backing repo: `@acme/*` from private, everything else from the cached public proxy. One line of config per machine, and the routing complexity lives in the registry where it belongs.

## Retention: artifacts pile up faster than you think

Every CI run can push an image. At a few builds an hour, you accumulate thousands of throwaway snapshot tags, and storage bills climb for builds nobody will ever pull again. Registries solve this with **retention (cleanup) policies** - rules for what to keep and what to reap.

```text
Retention policy (example shape):
  KEEP  release tags matching  v*.*.*           forever
  KEEP  the most recent        30  snapshot images
  DELETE  untagged images older than  14 days
```

*What just happened:* the policy keeps real releases indefinitely, keeps a rolling window of recent snapshots for debugging, and sweeps the untagged orphans that pile up when a moving tag is repointed (the old digest loses its tag but the bytes linger). Set this up early - retroactively cleaning a registry that's been hoarding for two years is a tense, careful job.

> Be specific about what "release" means in your keep rule. A too-broad delete rule that reaps an image production is still running is its own outage. Match release tags precisely, and prefer deleting *untagged* and *snapshot* artifacts over anything that looks like a version.

## Vulnerability scanning: know what you're shipping

An image isn't only your code - it's a base OS, system libraries, and every transitive dependency, any of which can carry a known CVE. Registries and CI tools scan artifacts against vulnerability databases so you find out *before* it's in production, not after.

```console
$ trivy image ghcr.io/acme/checkout-api:1.4.0
checkout-api:1.4.0 (debian 12.5)
Total: 4 (HIGH: 3, CRITICAL: 1)

┌────────────┬────────────────┬──────────┬───────────────┬──────────────┐
│  Library   │ Vulnerability  │ Severity │ Installed Ver │ Fixed Version│
├────────────┼────────────────┼──────────┼───────────────┼──────────────┤
│ libssl3    │ CVE-2024-XXXXX │ CRITICAL │ 3.0.11-1      │ 3.0.13-1     │
└────────────┴────────────────┴──────────┴───────────────┴──────────────┘
```

*What just happened:* the scanner cross-referenced every package in the image against CVE databases and flagged a critical one in a system library you never directly installed - it rode in on the base image. The "Fixed Version" column tells you the cure: rebuild on a patched base. Nexus and Artifactory run this server-side and can *block* a pull or promotion if an artifact is too risky; tools like Trivy or Grype run the same check in CI. The key caveat: a scan reflects what the database knew *at scan time*, so scanning continuously (not once at build) is what keeps it useful.

## The gotcha that bites everyone: the moving public tag

Here is the trap phase 1 promised. A tag is a movable pointer, and **most public registries let a tag be re-pushed to point at different content.** That convenience is a quiet hazard.

```console
# Monday - you build and deploy:
$ docker pull acme/widget:latest          # → sha256:aaaa  (the build you tested)

# Wednesday - someone re-pushes latest with a new build:
$ docker push acme/widget:latest          # latest now → sha256:bbbb

# Friday - autoscaler launches a new node, pulls "the same" image:
$ docker pull acme/widget:latest          # → sha256:bbbb  (different bytes!)
```

*What just happened:* nothing in your config changed, yet your fleet is now running two different builds - old nodes on `aaaa`, the new node on `bbbb` - because `latest` silently moved. You tested `aaaa`. Friday's node runs `bbbb`, untested, in production. This is the public-tag-overwrite bite: a tag you treated as a stable name was a pointer someone repointed.

Three defenses, strongest first:

```console
# 1. Deploy by digest - immutable, cannot be moved out from under you:
$ docker pull acme/widget@sha256:aaaa...

# 2. Use precise, never-reused version tags (1.4.0, never latest) in deploys.

# 3. Turn on immutable tags in your registry so a tag can't be re-pushed at all.
```

*What just happened:* pinning by digest (option 1) makes "what I tested" and "what runs" the same bytes by construction. Precise version tags (option 2) work if your team *disciplines* itself to never re-push them. Immutable-tag enforcement (option 3) removes the discipline requirement by making the registry reject any re-push of an existing tag - Artifactory, Nexus, GHCR, and the cloud registries all offer some form of this. Defense in depth: do all three for anything that reaches production.

## When to reach for which registry

Choose on purpose, not on hype:

- **Docker Hub / GHCR** when you want managed, zero-ops image hosting and you live in that ecosystem already. GHCR is the natural fit when your code is on GitHub - images sit beside repos under the same permissions.
- **Nexus / Artifactory** when you need *one* place for many formats (npm + Maven + PyPI + images), proxy/group repos, retention, server-side scanning, and fine-grained access - and you're willing to run (or pay for) the server. Artifactory leans enterprise-feature-rich; Nexus has a capable free tier. Both solve the same core problem.
- **Cloud registries (ECR/GCR/ACR)** when your workloads already run in that cloud - IAM integration and in-region pulls are the draw.

None is a moral choice. They all store and serve artifacts. The decision is formats, scale, and how much infrastructure you want to own.

## In the wild

A mature setup looks like this: one Artifactory or Nexus as the single front door, proxy repos in front of Docker Hub / npm / PyPI so CI never hits the public internet directly, group URLs so developers configure one endpoint per ecosystem, retention policies sweeping snapshot and untagged artifacts nightly, server-side scanning gating promotion to the `prod` repo, and production deploys pinned by digest. The registry is invisible until the day it saves you - the day Docker Hub rate-limits the world and your builds keep running because everything's already cached behind your own door.

```quiz
[
  {
    "q": "What problem does a proxy (remote) repository in Nexus or Artifactory primarily solve?",
    "choices": [
      "It encrypts your source code",
      "It caches artifacts pulled from a public registry, so later pulls are local and fast and survive public-registry rate limits or outages",
      "It automatically rewrites your Dockerfiles",
      "It deletes all untagged images"
    ],
    "answer": 1,
    "explain": "A proxy caches upstream pulls. After the first fetch, pulls are served locally - faster, and resilient to upstream rate limits and downtime. It's also a supply-chain control point."
  },
  {
    "q": "You deploy acme/widget:latest Monday (sha256:aaaa). Someone re-pushes latest Wednesday (sha256:bbbb). What happens when an autoscaler pulls latest Friday?",
    "choices": [
      "It fails because the tag changed",
      "It gets sha256:aaaa, the original build",
      "It gets sha256:bbbb - different, untested bytes than the rest of your fleet - because the tag was moved",
      "The registry blocks the pull automatically"
    ],
    "answer": 2,
    "explain": "A mutable tag was repointed, so the new node runs different bytes than the nodes you deployed Monday. This is the public-tag-overwrite trap; pin by digest or enforce immutable tags."
  },
  {
    "q": "Which is the strongest defense against a moving tag silently changing what production runs?",
    "choices": [
      "Always pull :latest so everyone is consistent",
      "Deploy by immutable digest (image@sha256:...), so the reference cannot be repointed",
      "Delete old images more often",
      "Increase the registry's storage quota"
    ],
    "answer": 1,
    "explain": "A digest is content-addressed and immutable - deploying by digest guarantees the bytes you tested are the bytes that run, no matter what happens to tags."
  }
]
```
