# Object Storage (S3)

> Buckets, keys, and signed URLs - how cloud object storage really works, what it is good and bad at, and the public-bucket leak that makes the news.


---

# Object Storage (S3)

You have files to store - user uploads, backups, a pile of images - and the disk on your server keeps filling up. Someone says "put it on S3," and you nod, but the moment you open the console it feels like a weird filesystem that lies to you: folders that aren't folders, permissions that bite, and a "Make public" button that has ended careers. This guide gives you the real mental model so object storage stops feeling like magic and starts feeling like a tool you trust.

## How to read this

Read the phases in order the first time. Phase 1 rewires how you picture the thing - and almost everything confusing about S3 comes from picturing it wrong. Phase 2 is the daily working knowledge. Phase 3 is where it bites people, including the leak you've read about in the news. If you only have five minutes, read phase 1; it's the part that saves you.

## The phases

1. [What it actually is: a giant key-to-blob map](01-the-mental-model.md)
2. [How you really work with it: keys, uploads, and signed URLs](02-keys-and-access.md)
3. [Where it bites: leaks, edits, and consistency](03-where-it-bites.md)


---

# What it actually is: a giant key-to-blob map

Here's the reality you arrived with: you opened the S3 console, saw something that *looks* like folders, double-clicked into one, and felt like you were browsing a hard drive in the sky. Then you tried to rename a "folder" and it wouldn't let you. Or you uploaded `report.pdf` and the URL had your whole path baked into it. That friction is your filesystem instinct rubbing against something that isn't a filesystem at all.

So let's replace the instinct.

## It's a dictionary, not a disk

Object storage is one enormous lookup table. You give it a **key** (a string), and it hands back a **blob** of bytes (the object). That's the whole model:

```text
KEY                                  ->  VALUE (the object's bytes + metadata)
-----------------------------------     --------------------------------------
"avatars/u-1843/profile.jpg"         ->  <the JPEG bytes, content-type, size...>
"backups/2026-06-30/db.sql.gz"       ->  <the gzip bytes...>
"invoices/2026/06/inv-90021.pdf"     ->  <the PDF bytes...>
```

*What just happened:* every object is found by its exact key and nothing else. There is no "open the avatars folder and look inside." There is only "fetch the object whose key is `avatars/u-1843/profile.jpg`." If you think of it as a Python dict or a JavaScript object - `store[key]` returns the bytes - you already understand the core.

## The slashes are a lie your eyes tell you

This is the single most important thing in the guide, so read it twice: **the `/` characters in a key are part of the key string. They are not directories.**

The key `invoices/2026/06/inv-90021.pdf` is one flat string. The storage system does not create a folder called `invoices`, then `2026` inside it, and so on. There is no tree. The console *renders* a fake folder view by splitting keys on `/` so your brain has something familiar to click - but underneath, it's a flat list of strings.

```text
What you see in the console        What actually exists (flat list of keys)
---------------------------        ----------------------------------------
📁 invoices/                       "invoices/2026/06/inv-90021.pdf"
   📁 2026/                        "invoices/2026/06/inv-90022.pdf"
      📁 06/                       "invoices/2026/07/inv-90023.pdf"
         📄 inv-90021.pdf
```

*What just happened:* the console grouped three flat strings by their `/` segments and drew folders, but no folder object exists anywhere. This is why you can't "rename a folder" - there's nothing to rename. To "move a folder," you copy every object to new keys and delete the old ones. There's no atomic directory rename because there's no directory.

> The technical term for the `/`-grouped view is a **prefix**. When a tool says "list objects with prefix `invoices/2026/`," it means "give me every key that starts with that string." Same idea as autocomplete, not folder navigation.

## A bucket is the namespace

A **bucket** is the top-level container that holds your keys. One AWS account can have many buckets; each bucket is its own keyspace and has its own name (globally unique across all of AWS, in the case of S3). So an object is fully identified by **bucket + key**:

```text
bucket:  acme-prod-uploads
key:     avatars/u-1843/profile.jpg
```

*What just happened:* together these two strings point to exactly one object anywhere in the cloud. The bucket is the "which dictionary," the key is the "which entry." That's the entire addressing scheme.

## Why anyone builds it this way

Throwing away the filesystem buys three things that matter enormously at scale:

- **It scales almost without limit.** There's no directory tree to lock, no inode table to grow, no single disk to fill. A key is a string in a distributed index, and the bytes get spread across many machines. You can store a handful of files or trillions; the model doesn't change.
- **It's cheap.** Because it's dumb and flat, the bytes can sit on commodity disks in bulk, and providers charge a small amount per gigabyte per month. Storing a terabyte costs roughly the price of a sandwich per month at standard tiers - far less than the equivalent always-on server disk.
- **It's extremely durable.** Each object is copied across multiple machines and often multiple buildings automatically. Providers quote durability like "eleven nines" (99.999999999%) for their standard class - a way of saying *they expect to almost never lose your object*. You don't manage the copies; that's the deal.

The price you pay for all that is the subject of phase 3: you give up in-place edits, instant directory operations, and the low latency of a local disk. For storing whole files you rarely change, that trade is a steal. For a database's hot files, it's a disaster.

## Where this fits

A regular server keeps your files on its own disk - fast, but it fills up and lives or dies with that one machine. (If "what a server even is" is fuzzy, see /guides/what-a-server-is.) Object storage is the cloud's answer to "I have a lot of files and I don't want to babysit disks." It's one of the foundational building blocks every cloud platform offers, alongside compute and databases - see /guides/cloud-platforms-explained for the bigger map.

For builders: the next time you reach for "save the uploaded file to `./uploads/`," pause. That works until you run a second server, or the box reboots, or the disk fills. Object storage is the standard home for user uploads precisely because it's not tied to any one machine.

```quiz
[
  {
    "q": "In a key like \"invoices/2026/inv-1.pdf\", what are the slashes?",
    "choices": [
      "Real directory separators that create nested folders",
      "Just characters inside one flat key string",
      "A required format that all keys must follow",
      "Pointers to other buckets"
    ],
    "answer": 1,
    "explain": "Keys are flat strings. The slashes are part of the string; the console only renders fake folders by grouping on them."
  },
  {
    "q": "What two things together uniquely identify an object?",
    "choices": [
      "The folder and the filename",
      "The region and the file extension",
      "The bucket and the key",
      "The account ID and the timestamp"
    ],
    "answer": 2,
    "explain": "An object is addressed by bucket (which keyspace) plus key (which entry in it)."
  },
  {
    "q": "Why can't you cheaply rename a 'folder' in object storage?",
    "choices": [
      "Renames require admin permissions you rarely have",
      "There is no folder - you must copy every object to new keys and delete the old ones",
      "The provider charges a large fee per rename",
      "Folder names are immutable by law"
    ],
    "answer": 1,
    "explain": "Folders don't exist. 'Renaming' a prefix means rewriting every object's key, since there's no directory to rename."
  }
]
```


---

# How you really work with it: keys, uploads, and signed URLs

Now you've got the model - bucket plus key, flat strings, no folders. This phase is the day-to-day: how the four operations work, how you choose good keys, and the one mechanism that solves the question every app eventually asks - *"how do I let this one user download this one private file without making it public to the whole internet?"*

## There are four verbs

Object storage gives you a tiny, blunt API. Once you see it, the rest is detail:

```text
PUT     bucket + key + bytes   ->  store (or fully overwrite) the object
GET     bucket + key           ->  fetch the object's bytes
DELETE  bucket + key           ->  remove the object
LIST    bucket + prefix        ->  return keys that start with that prefix
```

*What just happened:* that's the whole vocabulary. Notice what's missing - there's no APPEND, no "edit byte 500," no "rename." PUT writes the *entire* object every time. We'll come back to that limitation in phase 3; for now, just register that writes are whole-object.

Here's a real session with the AWS CLI so the verbs feel concrete:

```bash
# PUT: upload a local file to a key
aws s3 cp ./profile.jpg s3://acme-prod-uploads/avatars/u-1843/profile.jpg

# LIST: every key under a prefix (the fake "folder")
aws s3 ls s3://acme-prod-uploads/avatars/u-1843/

# GET: download it back
aws s3 cp s3://acme-prod-uploads/avatars/u-1843/profile.jpg ./got.jpg

# DELETE: remove it
aws s3 rm s3://acme-prod-uploads/avatars/u-1843/profile.jpg
```

*What just happened:* `s3 cp` is PUT or GET depending on direction, `s3 ls` is LIST scoped to a prefix, `s3 rm` is DELETE. The `s3://bucket/key` form is the address from phase 1 written as a URL.

## Choosing keys: this is your real "schema"

Because there are no folders to organize you, your **key naming convention is your data model.** Good prefixes make listing and lifecycle rules easy; bad ones make your life hard later. A pattern that holds up:

```text
<entity>/<id>/<purpose>/<filename>

users/1843/avatar/profile.jpg
orders/90021/invoice/inv-90021.pdf
backups/db/2026-06-30/full.sql.gz
```

*What just happened:* putting the stable, high-level grouping first (`users/`, `backups/db/`) means you can later say "list everything under `backups/db/2026-06-30/`" or "delete everything under `users/1843/`" with a single prefix. Date components in `YYYY-MM-DD` order sort correctly as plain strings, which makes "delete backups older than X" trivial.

> One caution: avoid putting a value that's identical across millions of objects at the *very front* of every key (like `uploads/<everything>`) if you're writing at extreme volume - historically that could concentrate load. For normal apps this never matters; mentioning it so the term "key prefix performance" isn't a mystery if you meet it.

## The access problem, and why "make it public" is the wrong reflex

Your app stores a user's private invoice at `orders/90021/invoice/inv-90021.pdf`. The user clicks "Download." How do you serve them the bytes?

The tempting answer is to flip the object (or worse, the whole bucket) to **public** so a plain URL works. Don't. Public means *the entire internet* can read it if they guess or discover the key - and keys are guessable (`inv-90021`, `inv-90022`...). That reflex is exactly what causes the leaks in phase 3.

The two correct patterns are:

- **Proxy it through your server.** Your app authenticates the user, fetches the object with its own credentials, and streams the bytes back. Simple and safe, but every byte flows through your server, which costs you bandwidth and CPU.
- **Hand out a signed URL.** Let the user download straight from the storage service, but only via a URL that's cryptographically stamped to expire. This is usually what you want.

## Signed URLs: a temporary, self-expiring key to one object

A **signed URL** (AWS calls it a *presigned URL*) is a normal-looking URL with extra query parameters: who's allowed, what action (GET or PUT), and crucially an **expiry**. The storage service checks the signature; if it's valid and unexpired, it serves the object. No login needed by the recipient - the URL *is* the credential.

```bash
# Generate a URL that lets the holder GET this one object for 15 minutes
aws s3 presign s3://acme-prod-uploads/orders/90021/invoice/inv-90021.pdf \
  --expires-in 900
```

It returns something shaped like this:

```text
https://acme-prod-uploads.s3.amazonaws.com/orders/90021/invoice/inv-90021.pdf
  ?X-Amz-Algorithm=AWS4-HMAC-SHA256
  &X-Amz-Credential=...
  &X-Amz-Date=20260630T120000Z
  &X-Amz-Expires=900
  &X-Amz-Signature=4a7c... (the cryptographic stamp)
```

*What just happened:* you generated, on your server (where your secret credentials live), a URL that grants exactly one action (GET) on exactly one object for exactly 900 seconds. You hand it to the authenticated user; their browser downloads directly from S3; fifteen minutes later the link is dead. The object stayed private the whole time - the bucket was never public.

The same trick works in reverse for **uploads**: generate a presigned PUT URL and the browser uploads the file straight to the bucket without the bytes ever touching your server. That's how big-file uploads avoid melting your app server.

```text
1. Browser asks your server: "I want to upload avatar.jpg"
2. Server (authenticated) generates a presigned PUT URL, scoped to one key, expiring soon
3. Browser PUTs the file directly to S3 using that URL
4. Browser tells your server "done"; server records the key in its database
```

*What just happened:* your server only handled a tiny signing request and a tiny confirmation. The heavy file transfer went browser-to-storage, which is faster for the user and cheaper for you. The URL's short expiry and single-key scope keep it safe even if it leaks.

For builders: think of a signed URL like a hotel key card. It opens one room, expires at checkout, and works without the front desk re-verifying who you are each time. You'd never make every room permanently unlocked (that's a public bucket) - you hand out a card that stops working soon.

```quiz
[
  {
    "q": "What does a PUT do to an object that already exists at that key?",
    "choices": [
      "Appends the new bytes to the end",
      "Fails with a conflict error",
      "Fully overwrites it with the new bytes",
      "Edits only the changed bytes in place"
    ],
    "answer": 2,
    "explain": "Writes are whole-object. PUT replaces the entire object; there's no append or in-place edit."
  },
  {
    "q": "What is the defining property of a signed (presigned) URL?",
    "choices": [
      "It makes the bucket public to everyone",
      "It grants a specific action on one object and expires after a set time",
      "It encrypts the object's contents",
      "It permanently authenticates the user's account"
    ],
    "answer": 1,
    "explain": "A signed URL is a scoped, time-limited credential: one action, one object, an expiry - no public bucket needed."
  },
  {
    "q": "Why use a presigned PUT URL for browser uploads?",
    "choices": [
      "It compresses the file automatically",
      "The file uploads directly to storage, never flowing through your app server",
      "It bypasses all permission checks",
      "It's the only way to upload files larger than 1 MB"
    ],
    "answer": 1,
    "explain": "The browser uploads straight to the bucket, so the heavy transfer skips your server - faster and cheaper, while the short-lived scoped URL stays safe."
  }
]
```


---

# Where it bites: leaks, edits, and consistency

You now know enough to use object storage well. This last phase is the part that keeps people out of the news and out of 3am pages: the famous public-bucket leak and how to never be it, the operations object storage is genuinely bad at, and the consistency gotcha that confused a generation of engineers. Plus the two things it's *perfect* for, so you reach for it at the right moments.

## The leak that makes the news

You've seen the headlines: "Company exposes millions of customer records in misconfigured S3 bucket." Almost every one of those is the same mistake. Someone needed one file reachable, flipped the bucket (or an object's access) to **public**, and forgot - or never understood - that "public" means *the entire internet can list and read it.*

Here's why it's so dangerous and so common:

```text
Public bucket "acme-uploads"
  ->  Anyone can do:  GET https://acme-uploads.s3.amazonaws.com/<any-key>
  ->  Anyone can do:  LIST the bucket and see EVERY key
```

*What just happened:* once a bucket is public, an attacker doesn't need to guess keys - they can list the whole thing and download everything. There's no login wall, no rate limit on curiosity. Researchers and bots scan for public buckets continuously; an exposed one is often found within hours.

The trap is that "public" feels like a reasonable answer to a reasonable question ("how do I serve this file?"). It is almost never the right answer for anything user-specific. The fix is the previous phase: keep buckets private, serve private files with **signed URLs**, and only ever make something public when it's *meant* for the whole world (your site's logo, public CSS).

> Modern S3 ships with **Block Public Access** turned on by default at the account and bucket level - a deliberate guardrail so you can't accidentally make a bucket public without consciously turning protections off. Treat that switch as a smoke detector: if you ever find yourself disabling it, stop and ask whether you really want the literal entire internet to have this data. The real answer is usually no.

The checklist that prevents the headline:

- **Default to private.** Leave Block Public Access on. Assume every object is sensitive until proven otherwise.
- **Serve private files via signed URLs**, not by toggling public.
- **Make public only what's genuinely public** - and even then, a dedicated bucket for public assets, separate from anything private, so a mistake can't expose customer data.
- **Don't rely on key obscurity.** A "secret" key like `exports/a8f3.../data.csv` is not access control; if the bucket is public, listing reveals it.

## What it's genuinely bad at

Object storage is a key-to-blob map, and that simplicity has real costs. Knowing them tells you when *not* to use it:

- **No in-place edits.** You cannot change byte 500 of an object. Every write is a full PUT of the whole object. Storing a frequently-mutated file (a database file, a log you append to constantly) means rewriting the entire thing on every change - wasteful and slow. Object storage is for *whole files you replace occasionally*, not data you poke at.
- **Higher latency than a local disk.** A GET is a network request to a remote service - tens of milliseconds, sometimes more, versus microseconds for local disk or RAM. Fine for serving an image or downloading a backup; a poor fit for anything in a tight loop that needs each byte *now*.
- **Listing is not a fast index.** LIST walks keys in lexical order with paging. It's great for "everything under this prefix," but it is not a query engine. You can't ask "all objects bigger than 1 MB modified last Tuesday." If you need to query *about* your files, keep that metadata in a real database and store only the bytes in the bucket.

The throughline: object storage is the wrong home for anything that wants to be edited in place, read with microsecond latency, or queried by attributes. It's the right home for the opposite - large, mostly-immutable blobs you fetch whole.

## The consistency gotcha (mostly history now)

For years, S3 was **eventually consistent** for some operations: you could write an object and a read a moment later might still return the *old* version (or a 404 for a brand-new key), because the change hadn't propagated to every replica yet. This bit people hard - "I uploaded it, why does my code say it's not there?"

```text
T+0ms   PUT  invoices/new.pdf        (write accepted)
T+5ms   GET  invoices/new.pdf        -> 404 Not Found   (replica hadn't caught up)
T+50ms  GET  invoices/new.pdf        -> 200 OK          (now it's there)
```

*What just happened:* the write succeeded, but a read a few milliseconds later hit a replica that didn't have it yet. The data wasn't lost - the system just hadn't finished agreeing with itself.

The good news: **S3 now provides strong read-after-write consistency** - a successful write is immediately visible to a following read. You generally don't have to engineer around this on S3 anymore. But keep the concept in your pocket, because (a) plenty of older code still has sleeps and retries built to dodge it, and (b) other object stores and distributed systems may still be eventually consistent. When a freshly-written object "isn't there yet," eventual consistency is the first suspect on systems that have it.

## What it's perfect for

End on the bright side - the two jobs object storage was born to do:

- **Static hosting.** Your site's images, CSS, JS, downloads - large, immutable, requested by everyone. Object storage serves them cheaply and durably, usually with a CDN in front for speed. These are the *legitimately public* objects; a dedicated public bucket is exactly right here.
- **Backups and archives.** Database dumps, log archives, "we might need this in three years" data. Cheap per gigabyte, eleven-nines durable, and you can set lifecycle rules to auto-delete old backups or move cold data to even cheaper tiers. This is the canonical use: write once, read rarely, keep forever, pay little.

For builders: a clean default architecture is *private bucket + signed URLs for user files, separate public bucket + CDN for static assets, lifecycle rules for backups.* That covers the vast majority of real apps and keeps you off the leak list. If "where does the bucket live and who runs it" is still hazy, /guides/cloud-platforms-explained puts object storage next to the other cloud building blocks.

```quiz
[
  {
    "q": "What is the root cause of the classic 'company leaked data in S3' headline?",
    "choices": [
      "A bug in S3's encryption",
      "A bucket or object set to public, exposing it to the entire internet",
      "Hackers brute-forcing AWS passwords",
      "Signed URLs that never expired"
    ],
    "answer": 1,
    "explain": "Public buckets let anyone list and download everything. The fix is to stay private and use signed URLs for private files."
  },
  {
    "q": "Which workload is object storage a POOR fit for?",
    "choices": [
      "Serving your website's images and CSS",
      "Storing nightly database backups",
      "A database file that's edited in place many times per second",
      "Holding user-uploaded photos"
    ],
    "answer": 2,
    "explain": "There are no in-place edits - every write is a full PUT. Frequently-mutated files are exactly what object storage is bad at."
  },
  {
    "q": "On modern S3, what consistency do you get after a successful write?",
    "choices": [
      "Strong read-after-write consistency - a following read sees the new data immediately",
      "Eventual consistency, so you must sleep and retry",
      "No consistency guarantees at all",
      "Consistency only if you pay for a premium tier"
    ],
    "answer": 0,
    "explain": "S3 now provides strong read-after-write consistency. Older code may still have retries from the eventually-consistent days, and other stores may still be eventual."
  }
]
```
