# Protobuf and Avro

> Binary serialization with a schema: Protocol Buffers and Avro make data small and fast across services, and force you to think about schema evolution.


---

# Protobuf and Avro

You have been shipping JSON between services for years and it works fine, until the day the payloads get huge, the parsing gets slow, or someone renames a field in a producer and a dozen consumers fall over with no warning. JSON is plain and readable, but it carries its field names in every single message and trusts everyone to agree on the shape by hand. Protobuf and Avro fix both problems: a real schema, a compact binary wire format, and rules for changing that schema without breaking the world. This guide gives you the mental model so the two stop blurring together, then the everyday mechanics, then the part nobody warns you about: evolution.

## How to read this

Read the phases in order the first time. Phase 1 is the why, and skipping it is how people end up cargo-culting `.proto` files they do not understand. Phase 2 is the hands-on core you will reach for daily. Phase 3 is the production reality, schema evolution and the traps, and it is the reason these tools exist at all, so do not stop early.

## The phases

1. [The mental model: schema-first binary data](01-why-schemas-beat-json.md) - why a schema plus binary beats schema-less JSON, and how Protobuf and Avro split the problem.
2. [Using them day to day](02-using-protobuf-and-avro.md) - writing a `.proto` and an Avro schema, generating code, and reading the wire.
3. [Schema evolution and the gotchas](03-schema-evolution-and-gotchas.md) - forward and backward compatibility, the field-number rules, and where it bites in production.


---

# The mental model: schema-first binary data

Here is the thing JSON never tells you: every message you send carries its own field names, as plain text, every single time. Send a million records and you send the string `"customer_id"` a million times. That is fine when a human is reading a config file. It is a tax when machines are talking to machines at volume.

So picture the actual cost. Here is one record on the wire as JSON:

```json
{"customer_id": 4291, "is_active": true, "balance_cents": 19900}
```

*What just happened:* that is 61 bytes, and roughly 40 of them are field names and punctuation, not data. The receiver also has to parse text, guess that `4291` is a number and not a string, and trust that the sender used exactly the keys it expected. Nobody enforced any of that.

## What a schema actually buys you

A schema is a contract written down in one place: these are the fields, these are their types, this is their order. Both sides agree on it ahead of time. Once they do, two things become possible that JSON cannot do.

First, the wire format gets tiny. If both sides already know field 1 is `customer_id` and it is an integer, the message does not need to say `"customer_id"` at all. It can say "field 1, value 4291" using a couple of bytes. The names live in the schema, not in every message.

Second, the data gets typed and checked. The schema says `balance_cents` is a 64-bit integer, so the generated code gives you an actual integer, not a string you have to coerce and pray over. A producer that tries to put a name where a number goes fails at build time, not at 3am in a consumer's logs.

> The trade is real and worth naming. JSON you can read with your eyes in any text editor. A Protobuf or Avro payload is opaque bytes; you need the schema to decode it. You are trading human readability for size, speed, and a contract. For machine-to-machine traffic at scale, that trade is usually worth it. For a config file a person edits, it is not.

## Protobuf and Avro split the problem differently

Both are schema-first binary formats, and people lump them together, but they answer one key question in opposite ways: **where does the schema live when the data is decoded?**

**Protocol Buffers (Protobuf)**, from Google, is *code-generation first*. You write a `.proto` file, run a compiler (`protoc`), and it generates classes in your language. The schema is baked into the compiled code on both sides. The wire format uses small integer *field numbers* instead of names, which is why it is compact and fast. Protobuf is the serialization backbone of gRPC, so if you are doing service-to-service RPC, this is the one you will meet. For the RPC side of that story, see [/guides/grpc-explained](/guides/grpc-explained).

**Avro**, from the Hadoop world, is *schema-travels-with-data first*. The schema is a JSON document, and the classic Avro file format writes the full schema once into the header of the file, then packs thousands of records after it with no per-record overhead. Because the reader can read the schema out of the file, it does not need generated code at all, it can decode dynamically. This is why Avro dominates the data and streaming world: Hadoop, Spark, and especially Kafka, where huge volumes of records share one schema.

```text
Protobuf:  schema (.proto) --compile--> code on both sides
           wire = field numbers + values   (names never sent)

Avro:      schema (JSON) lives WITH the data (file header / registry)
           wire = values in schema order    (names never sent)
```

*What just happened:* both drop field names from the wire, but Protobuf assumes both sides already hold compiled code, while Avro assumes the reader can get the schema from the data or a registry. That single difference drives almost everything else, including how each handles change.

## The real reason they matter: change is coming

Small and fast is the headline, but it is not why teams adopt these tools and stay. The deeper reason is **schema evolution**. Your data shape will change, you will add a field, drop one, rename something, and in a distributed system you cannot redeploy every producer and consumer at the same instant. For a while, old code and new code run side by side, reading each other's messages.

JSON gives you nothing here. There is no schema, so there is no notion of "is this change safe?" You find out in production. Protobuf and Avro both have explicit rules, and tooling to check them, for whether a change is *backward compatible* (new code reads old data) and *forward compatible* (old code reads new data). That is the actual superpower, and it gets its own phase, because it is where most of the real-world pain and most of the real-world value live.

## In the wild

If you have ever called a gRPC service, you have used Protobuf without writing a line of it, the framing was Protobuf under the hood. If you have ever consumed a Kafka topic backed by a Schema Registry, you have used Avro the same way. Most engineers meet these formats as a dependency of something bigger before they ever write a schema by hand.

```quiz
[
  {
    "q": "What is the main thing a schema lets the wire format drop, making it compact?",
    "choices": ["Numeric values", "Field names", "The message length", "Timestamps"],
    "answer": 1,
    "explain": "Both sides know the fields from the schema, so the bytes carry values keyed by small field numbers or position, not repeated field-name strings."
  },
  {
    "q": "What best describes the core difference between Protobuf and Avro?",
    "choices": ["Protobuf is binary, Avro is text", "Protobuf generates code from a .proto; Avro often ships the schema with the data and can decode dynamically", "Avro is faster on every workload", "Protobuf has no schema"],
    "answer": 1,
    "explain": "Protobuf is code-generation-first with compiled schemas on both sides; Avro carries its schema with the data (file header or registry), so it can decode without generated code."
  },
  {
    "q": "Beyond small size and speed, what is the deeper reason teams adopt these formats?",
    "choices": ["They are human-readable", "They need no schema", "Schema evolution: explicit rules for changing data shape without breaking running producers and consumers", "They replace databases"],
    "answer": 2,
    "explain": "In a distributed system you cannot redeploy everything at once; Protobuf and Avro define forward/backward compatibility rules so old and new code interoperate during rollout."
  }
]
```


---

# Using them day to day

The mental model lands fastest when you write the schemas yourself. So let us do that, the same little `Customer` record in both formats, generate code, and look at what actually goes on the wire. The workflows feel different, and the difference tells you a lot about each tool's personality.

## Protobuf: write the schema, compile it, use the classes

A Protobuf schema is a `.proto` file. Here is a small one.

```proto
syntax = "proto3";

package billing;

message Customer {
  int64  id           = 1;
  string name         = 2;
  bool   is_active    = 3;
  int64  balance_cents = 4;
}
```

*What just happened:* you declared a message type and, crucially, gave every field a *number* (`= 1`, `= 2`, …). Those numbers, not the names, are what travel on the wire. The names are for humans and generated code. Hold on to this fact: in Phase 3 those field numbers turn out to be the load-bearing part of the whole compatibility story.

Now compile it. The `protoc` compiler reads the `.proto` and emits code for your target language.

```bash
# generate Python classes into the ./gen directory
protoc --python_out=gen customer.proto

# or Go
protoc --go_out=gen customer.proto
```

*What just happened:* `protoc` produced a generated module (for Python, `customer_pb2.py`) holding a real `Customer` class with typed fields, plus serialize and parse methods. You do not hand-write parsing code; you use the generated class.

Then you use it like any object, and serialize to bytes when you need to send.

```python
from gen import customer_pb2

c = customer_pb2.Customer(id=4291, name="Mara", is_active=True, balance_cents=19900)

data = c.SerializeToString()   # -> compact binary bytes
print(len(data))               # far smaller than the JSON equivalent

# on the other side, with the same generated class:
c2 = customer_pb2.Customer()
c2.ParseFromString(data)
print(c2.name)                 # "Mara"
```

*What just happened:* the sender turned a typed object into a few dozen bytes, and the receiver turned those bytes back into a typed object, both using code generated from the same `.proto`. Neither side ever sent or parsed the string `"name"`. This is the Protobuf loop you will live in: edit `.proto`, run `protoc`, use the classes.

> Field names are absent from the wire, so renaming `name` to `full_name` in the `.proto` and recompiling both sides changes nothing on the wire, because field 2 is still field 2. That is a feature for compatibility and a footgun if you misread it. More on that next phase.

## Avro: a JSON schema, and the data carries it

An Avro schema is itself a JSON document. The same record looks like this.

```json
{
  "type": "record",
  "name": "Customer",
  "namespace": "billing",
  "fields": [
    {"name": "id",            "type": "long"},
    {"name": "name",          "type": "string"},
    {"name": "is_active",     "type": "boolean"},
    {"name": "balance_cents", "type": "long"}
  ]
}
```

*What just happened:* you described the same four fields, but notice there are no field numbers. Avro identifies fields by *name and position in the schema*, not by an integer tag. That is a different identity model from Protobuf, and it changes the evolution rules you will meet in Phase 3.

Avro's signature move is the *object container file*: it writes the full schema once into the file header, then packs records after it. Because the schema is right there in the file, a reader needs no generated code, it reads the schema, then decodes the records.

```python
import fastavro, io

schema = fastavro.parse_schema({
    "type": "record", "name": "Customer", "namespace": "billing",
    "fields": [
        {"name": "id", "type": "long"},
        {"name": "name", "type": "string"},
        {"name": "is_active", "type": "boolean"},
        {"name": "balance_cents", "type": "long"},
    ],
})

records = [{"id": 4291, "name": "Mara", "is_active": True, "balance_cents": 19900}]

buf = io.BytesIO()
fastavro.writer(buf, schema, records)     # header carries the schema, then the rows

buf.seek(0)
for rec in fastavro.reader(buf):          # reader pulls the schema out of the header
    print(rec["name"])                    # "Mara"
```

*What just happened:* the writer stamped the schema into the file header, then wrote the records with no per-row field names. The reader recovered the schema from the header and decoded the rows, no generated `Customer` class required. This is exactly why Avro fits batch and streaming systems: write the schema once, stream a billion rows behind it.

## The Kafka twist: a Schema Registry, not a file header

In streaming, you are not writing one big file; you are sending many small messages to a topic. Stamping the whole schema into every Kafka message would undo the savings. So the ecosystem uses a **Schema Registry**: a small service that stores schemas and hands each one an integer ID.

```text
Producer:  register schema -> get id 7
           message on wire = [magic byte][id=7][Avro-encoded value]

Consumer:  read id 7 from the message
           fetch schema #7 from the registry (and cache it)
           decode the value
```

*What just happened:* each Kafka message carries a tiny schema ID instead of the whole schema, and consumers look the ID up once and cache it. You get Avro's "schema travels with data" guarantee without paying for the schema in every message. This registry pattern is the standard way Avro and Kafka work together, and the registry is also where compatibility gets enforced, which is the heart of the next phase.

## For builders

Pick by where the data is going. Talking RPC between services? Reach for Protobuf, because gRPC expects it and the codegen workflow fits request/response shapes, see [/guides/grpc-explained](/guides/grpc-explained). Moving events or analytics records through Kafka or a data lake? Reach for Avro, because the registry workflow and dynamic decoding fit high-volume streams. Plenty of shops run both, each in its lane. And if you are still mapping how services talk at all, [/guides/what-an-api-is](/guides/what-an-api-is) is the ground floor under this.

```quiz
[
  {
    "q": "In a .proto file, what do the numbers like `= 1` and `= 2` after each field represent?",
    "choices": ["Default values", "The field's position on screen", "Field tags that identify the field on the wire instead of its name", "The maximum value the field can hold"],
    "answer": 2,
    "explain": "Protobuf sends those integer field numbers, not field names. The names exist only in the schema and generated code."
  },
  {
    "q": "How does the classic Avro object container file let a reader decode without generated code?",
    "choices": ["It includes the field names in every record", "It writes the full schema into the file header, so the reader reads the schema then the records", "It calls protoc at read time", "It stores data as JSON text"],
    "answer": 1,
    "explain": "Avro stamps the schema once into the file header; the reader recovers it from there and decodes the records dynamically."
  },
  {
    "q": "Why do Kafka + Avro setups use a Schema Registry instead of putting the schema in every message?",
    "choices": ["Kafka cannot store binary data", "Embedding the full schema in each small message would erase the size savings; a registry lets each message carry a tiny schema ID instead", "Avro requires JSON over the wire", "Registries make messages human-readable"],
    "answer": 1,
    "explain": "Each message carries a small integer schema ID; consumers fetch and cache the schema once from the registry, keeping messages compact."
  }
]
```


---

# Schema evolution and the gotchas

This is the phase that earns the whole guide. Small and fast got your attention; schema evolution is what keeps a distributed system from melting down every time someone adds a field. In production, producers and consumers do not upgrade in lockstep. For minutes or days, new code and old code run side by side, reading each other's bytes. The schema's job is to make that overlap safe, and to tell you, before you ship, when a change is not.

## Two words you have to keep straight

Everything here is one of two questions. Say them out loud when you review a change:

- **Backward compatible:** can **new** code read **old** data? You upgraded the reader first; it must still understand messages written by the old writer.
- **Forward compatible:** can **old** code read **new** data? You upgraded the writer first; the not-yet-upgraded reader must not choke on the new messages.

```text
                 reads old data        reads new data
new reader   ->   BACKWARD                 (n/a)
old reader   ->     (n/a)                FORWARD
```

*What just happened:* the direction is named after the data, not the code. "Backward compatible" means the new thing reaches *back* to old data. Full compatibility means a change is both, which is what you usually want for an independently deployed system.

## Protobuf: the rules live in the field numbers

Remember those field numbers from the `.proto`. They are the identity of each field, and the evolution rules fall right out of that.

```proto
message Customer {
  int64  id            = 1;
  string name          = 2;
  bool   is_active      = 3;
  int64  balance_cents = 4;
  string email          = 5;   // newly added field
}
```

*What just happened:* you added `email` with a **new, never-before-used number 5**. An old reader that does not know field 5 sees those bytes, does not recognize the tag, and skips them, so old code reads new data fine (forward compatible). A new reader handed old data finds field 5 absent and uses the default (empty string), so new code reads old data fine (backward compatible). Adding an optional field with a fresh number is safe in both directions.

The rules that keep Protobuf safe, and the ones that wreck it:

- **Never reuse or change a field number.** The number is the identity. If you delete field 4 and later add a different field as number 4, old data's `balance_cents` bytes get read as the new field, silent corruption. Use `reserved 4;` to fence off retired numbers so nobody reuses them.
- **Renaming a field is free on the wire.** `name` to `full_name` keeps number 2, so the bytes are identical. Only your source code changes.
- **Do not change a field's type incompatibly.** Swapping `int64` to `string` on the same number reinterprets the bytes and breaks readers. Some numeric widenings are compatible; arbitrary type swaps are not.

```proto
message Customer {
  reserved 4;              // balance_cents retired; never reuse this number
  reserved "balance_cents";
  int64  id   = 1;
  string name = 2;
}
```

*What just happened:* `reserved` tells the compiler to reject any future field that tries to claim number 4 or the old name. It is a tombstone, and it is the difference between a safe deletion and a time bomb.

## Avro: the rules live in defaults and the reader/writer schema pair

Avro has no field numbers. It matches fields by **name**, and it does something Protobuf does not: at decode time it uses **two** schemas, the *writer's* schema (what made the data) and the *reader's* schema (what your code wants). Avro resolves the difference between them. That pairing is the core of Avro evolution.

- **To add a field, give it a `default`.** When old data (writer's schema lacks the field) is read by new code (reader's schema has it), Avro fills in the default. No default, and reading old data fails.
- **To remove a field, the field you drop must have had a default** so that new data (missing it) can still be read by old code expecting it.
- **Renaming uses `aliases`.** Because Avro matches on name, a raw rename looks like "old field gone, new field added." An `alias` tells the reader the new name also answers to the old one.

```json
{
  "type": "record", "name": "Customer",
  "fields": [
    {"name": "id",        "type": "long"},
    {"name": "name",      "type": "string"},
    {"name": "is_active", "type": "boolean"},
    {"name": "email",     "type": "string", "default": ""}
  ]
}
```

*What just happened:* `email` carries `"default": ""`. A reader on this schema decoding older data that has no `email` substitutes the default instead of failing. That single `default` is what makes the add backward compatible. In Avro, defaults are not a convenience, they are the evolution mechanism.

## The Schema Registry enforces this for you

You do not have to remember all of that by hand under deadline. A Schema Registry (the Kafka pattern from Phase 2) checks every new schema version against a configured **compatibility mode** before it accepts it.

```text
BACKWARD (the common default): new schema can read data written by the previous schema
FORWARD:  previous schema can read data written by the new schema
FULL:     both directions
NONE:     no checks (you are on your own)
```

*What just happened:* register an incompatible schema under `BACKWARD` mode and the registry **rejects it at registration**, before a single bad message is produced. The compatibility rule becomes a gate in your pipeline, not a postmortem. Set the mode to match how you deploy: upgrade consumers first, lean `BACKWARD`; upgrade producers first, lean `FORWARD`; want freedom in either order, use `FULL`.

## Gotchas that bite real teams

> **Reusing a Protobuf field number is the classic disaster.** It does not error. Old bytes get reinterpreted as the new field and you ship corrupted reads that pass every type check. Always `reserved` deleted numbers, and treat the number space as append-only forever.

A few more that show up in incident reviews:

- **proto3 and the meaning of "missing."** In plain proto3 scalar fields, a field set to its default (`0`, `""`, `false`) is indistinguishable on the wire from a field never set. If "absent" must differ from "zero," model it explicitly (for example with an `optional` field or a wrapper) rather than assuming you can tell them apart.
- **Required fields are a trap.** proto2 had `required`, and it made evolution nearly impossible, you can never safely remove a required field. proto3 dropped it on purpose. Resist any urge to simulate hard-required fields at the schema layer; enforce required-ness in application logic.
- **Compatibility is transitive, or it should be.** Checking each new version only against the immediately previous one lets data drift across many hops. If you keep long histories of data around, use the registry's *transitive* modes so a new schema is checked against all prior versions, not only the last.
- **The schema and the data must not drift apart.** Lose the `.proto` that compiled a producer, or fail to register an Avro schema, and you have opaque bytes nobody can decode. Version schemas in source control; treat the registry as production infrastructure, with backups.

## In the wild

The teams who stay calm during rollouts are the ones who made compatibility a build-time gate: schemas in source control, a registry in `FULL` or `BACKWARD` mode, CI that fails the pipeline when a `.proto` or Avro change would break a live consumer. The format is not what saves you, the discipline of checking evolution before you ship is. Protobuf and Avro give you a place to enforce it.

```quiz
[
  {
    "q": "A change is described as 'backward compatible.' What does that guarantee?",
    "choices": ["Old code can read new data", "New code can read old data", "The wire format shrinks", "Field names are preserved"],
    "answer": 1,
    "explain": "Backward compatibility means the new reader reaches back to data written by the old writer: new code reads old data."
  },
  {
    "q": "In Protobuf, why is reusing a deleted field number so dangerous?",
    "choices": ["It throws a compile error every time", "Old data's bytes for that number get silently reinterpreted as the new field, corrupting reads with no error", "It doubles the message size", "It deletes the registry"],
    "answer": 1,
    "explain": "The number is the field's identity on the wire. Reusing it makes old bytes decode as the new field silently. Mark deleted numbers `reserved`."
  },
  {
    "q": "In Avro, what makes adding a new field backward compatible so readers can decode older data that lacks it?",
    "choices": ["Assigning it a field number", "Giving the field a `default` value", "Marking it `required`", "Putting it first in the record"],
    "answer": 1,
    "explain": "Avro matches by name and resolves reader vs writer schemas; a `default` lets the reader fill in the value when older data omits the field."
  }
]
```
