# Character Encodings and Unicode

> Why text turns into garbled symbols, how UTF-8 actually works, and the difference between bytes, code points, and the characters a user sees.


---

# Character Encodings and Unicode

You opened a file and where a name should be, you got `Ã©` or `â€™` or a row of black diamonds with question marks. Or your code said a string was 7 characters long when the user clearly typed 5. Text feels like it should be the simplest thing a computer handles, and yet it betrays you constantly. The reason is almost always the same: somewhere, bytes were treated as characters, and they are not the same thing.

This guide gives you the one mental model that makes all of it click, then walks the real machinery, then shows you exactly where it breaks and how to never get burned again.

## How to read this

Read the phases in order the first time. Phase 1 installs the core distinction (bytes are not characters) that every later confusion depends on. Phase 2 is the everyday mechanics of UTF-8 and where most working programmers spend their time. Phase 3 is the deep end: emoji, combining characters, and why `length` lies. If you only have five minutes, read Phase 1 and the first section of Phase 2.

## The phases

1. [Bytes Are Not Characters](01-bytes-are-not-characters.md)
2. [How UTF-8 Actually Works](02-how-utf-8-actually-works.md)
3. [When Text Lies: Emoji, Graphemes, and the BOM](03-when-text-lies.md)


---

# Bytes Are Not Characters

Here's the thing nobody tells you up front, and it's the source of nearly every text bug you'll ever hit: **a computer never stores characters. It stores numbers.** A file on disk, a packet on the wire, a string in memory - all of it is bytes, which are nothing more than numbers from 0 to 255. The letter `A` is not in that file. A *number* is in that file, and somewhere there is an agreement that says "when you see this number, draw an `A`."

That agreement is called an **encoding**. The whole topic of character encoding is the story of who agreed on what, what happens when two parties disagree, and how the world slowly built one table big enough to hold every character humans use.

## The two-step dance: encode and decode

Every piece of text makes a round trip. When you save text, the computer **encodes** it: it takes the characters in your head and turns them into bytes. When you open text, the computer **decodes** it: it takes bytes off the disk and turns them back into characters to show you.

```text
"Hi"  --encode-->  [72, 105]  --decode-->  "Hi"
characters          bytes               characters
```

*What just happened:* The string `Hi` became two numbers, 72 and 105, and then came back. The bytes in the middle are the only thing that ever touched the disk. The characters at either end only exist while a program is holding them.

The catch is the decode step. To turn bytes back into characters, the decoder has to know **which agreement was used to make them**. If you encode with one agreement and decode with a different one, you get the wrong characters out. That mismatch has a name, and you have seen it.

## Where it all started: ASCII

The original agreement was **ASCII**. It is a tiny table - 128 entries - mapping numbers 0 through 127 to the characters an American typewriter cared about: the uppercase and lowercase Latin letters, the digits, basic punctuation, and a handful of control codes like newline and tab.

```text
65 -> A      97  -> a      48 -> 0
66 -> B      98  -> b      32 -> (space)
67 -> C      99  -> c      10 -> (newline)
```

*What just happened:* These are the actual ASCII numbers. `A` is 65, `a` is 97 (lowercase letters sit 32 above their uppercase twins, which is why `c - 'a'` math works). Every plain-text file made of English fits in this table, one byte per character.

ASCII had one beautiful property: every character fit in a single byte, with room to spare (a byte holds 0–255; ASCII only used 0–127). For English, the world was simple. One byte, one character, no ambiguity. And then everyone who did not speak English showed up.

## The chaos years: code pages

A byte can hold 256 values, but ASCII only claimed the first 128. That left 128 unused slots (128–255), and every region on earth filled them with *their own* characters. Western Europe put accented letters there. Greece put the Greek alphabet there. Russia put Cyrillic there. These regional tables were called **code pages**, and there were dozens of them, all conflicting.

The disaster was inevitable. The same byte meant different things in different places:

```text
byte 233 decoded as Latin-1 (Western Europe)  ->  é
byte 233 decoded as Windows-1251 (Cyrillic)   ->  щ
byte 233 decoded as Mac Roman                 ->  È
```

*What just happened:* One identical byte, 233, produces three completely different characters depending on which code page the decoder assumes. The byte carries no label saying which table it belongs to. The reader has to guess - and when it guesses wrong, you get garbage.

> A byte does not know its own encoding. Nothing inside the number 233 tells you it means `é`. The encoding is context the reader supplies, and if that context is wrong, the text is wrong.

## Mojibake: the wrong-decoder bug

When bytes are decoded with the wrong agreement, the result is **mojibake** (a Japanese word, roughly "character transformation"). It is the `Ã©` and `â€™` you have seen in badly handled text. It is not random corruption - the bytes are perfectly intact. They are being *read* through the wrong table.

```text
Author writes:   café
Saved as UTF-8:  [99, 97, 102, 195, 169]
Read as Latin-1: c  a  f  Ã   ©
You see:         cafÃ©
```

*What just happened:* The é was encoded as two bytes in UTF-8 (195 and 169 - Phase 2 explains why two). When a reader decodes those same two bytes one at a time using Latin-1, byte 195 draws `Ã` and byte 169 draws `©`. The data is fine. The interpretation is broken. Re-decode the *same bytes* with UTF-8 and `café` comes right back.

This is the single most important takeaway of the whole guide: **mojibake is a decode-side mismatch, not data loss.** When you see it, the fix is almost never "the file is corrupt." The fix is "tell the reader the correct encoding."

For builders: this is why "what encoding?" is the first question to ask whenever text looks wrong. Not "is the file broken" - the file is usually fine. The reader was handed the wrong table.

```quiz
[
  {
    "q": "A text file contains the byte 233. What character does it represent?",
    "choices": ["é, always", "It depends on which encoding you decode it with", "A question mark", "Nothing - 233 is invalid"],
    "answer": 1,
    "explain": "A byte carries no encoding label. The same byte 233 is é in Latin-1, щ in Windows-1251, and so on. The decoder's chosen table decides."
  },
  {
    "q": "You open a file and see cafÃ© where café should be. What most likely happened?",
    "choices": ["The file is corrupted and data was lost", "The bytes were encoded as UTF-8 but decoded as Latin-1", "ASCII cannot store the letter c", "The disk has a hardware fault"],
    "answer": 1,
    "explain": "This is mojibake: a wrong-decoder mismatch. The bytes are intact; re-decoding the same bytes with the correct encoding (UTF-8) restores café."
  },
  {
    "q": "Why did code pages cause so many problems?",
    "choices": ["They were too slow", "Each region reused the same byte values (128-255) for different characters, with no label saying which", "They could not store English", "They used two bytes per character"],
    "answer": 1,
    "explain": "Code pages all claimed the upper 128 byte values for different characters. Identical bytes meant different things in different places, and nothing in the data said which."
  }
]
```


---

# How UTF-8 Actually Works

The code-page chaos from Phase 1 had an obvious cure: stop having dozens of conflicting 256-entry tables and build **one table for every character on earth**. That table is **Unicode**. But Unicode by itself only solves half the problem - it assigns numbers; it does not say how to store them in bytes. The part that actually runs the modern internet is the *byte* encoding called **UTF-8**. This phase pulls those two apart, because confusing them is where intermediate developers get stuck.

## Unicode is the catalogue, not the storage

Unicode does one job: it gives every character a unique number called a **code point**. That is it. `A` is code point 65. `é` is code point 233. The euro sign `€` is 8364. The grinning face emoji is 128512. Code points are written in a standard hex form like `U+0041` (that is 65, the letter `A`).

```text
A    ->  U+0041  (65)
é    ->  U+00E9  (233)
€    ->  U+20AC  (8364)
😀   ->  U+1F600 (128512)
```

*What just happened:* These are the canonical Unicode code points. Notice they got large - the emoji is over a hundred thousand. A single byte stops at 255, so there is no way to cram these numbers into one byte each. Unicode the catalogue says *what number* a character is. It pointedly does **not** say how to write that number as bytes on disk. That is a separate decision, and there is more than one way to make it.

The cleanest mental model: **Unicode = the dictionary of characters and their numbers. The encoding = how you serialize those numbers into bytes.** Same dictionary, several possible byte formats.

## UTF-8: the encoding that won

The encoding the world settled on is **UTF-8**. It is a *variable-width* encoding: a character takes between one and four bytes depending on how big its code point is. This is its superpower. It packs the common case tight and only spends extra bytes when it has to.

```text
code point range        bytes used
U+0000  – U+007F        1 byte
U+0080  – U+07FF        2 bytes
U+0800  – U+FFFF        3 bytes
U+10000 – U+10FFFF      4 bytes
```

*What just happened:* The smaller the code point, the fewer bytes. Every ASCII character (U+0000–U+007F) is exactly one byte in UTF-8 - and it is the *same* byte ASCII always used. That is the killer feature: **valid ASCII is already valid UTF-8.** Decades of English text and existing code "just work" with zero conversion. The accented `é` and the euro sign cost two or three bytes; emoji cost four.

So a string's byte length and its character count are no longer the same number - and that gap is exactly what trips people up.

```text
"héllo"   ->  5 characters, but 6 bytes (the é takes 2)
"€5"      ->  2 characters, but 4 bytes (the € takes 3)
```

*What just happened:* `héllo` looks like five characters because it is - but on disk it is six bytes, because the `é` quietly costs two. Any code that assumes "bytes = characters" (slicing a string at byte 3, for example) can cut a multi-byte character in half and produce garbage.

## Seeing the bytes for real

Here is the round trip from Phase 1, now made concrete. You can run this.

```python runnable
s = "café"
encoded = s.encode("utf-8")
print("characters:", len(s))
print("bytes:", len(encoded))
print("byte values:", list(encoded))
print("round trip:", encoded.decode("utf-8"))
```

*What just happened:* The string has 4 characters but 5 bytes - the `é` is the two bytes `[195, 169]`, exactly the pair from Phase 1's mojibake example. `encode("utf-8")` turned characters into bytes; `decode("utf-8")` turned them back. Decode those same bytes with `"latin-1"` instead and you would get `café` mangled into `cafÃ©`.

## Why UTF-8 beat the alternatives

There were other ways to encode Unicode. UTF-16 uses two bytes for most characters; UTF-32 uses a flat four bytes for everything. Both waste space on English-heavy text and, worse, both raise a nasty question: when a character is more than one byte, **which byte comes first?**

```text
The same code point U+0041 (A) in UTF-16:
big-endian:    [0x00, 0x41]
little-endian: [0x41, 0x00]
```

*What just happened:* Multi-byte encodings have to decide byte order ("endianness"), and the two orders are incompatible. Get it wrong and every character is scrambled. UTF-16 and UTF-32 carry this hazard. UTF-8 sidesteps it entirely - its byte order is fixed by the format itself, so there is never an endianness question. That, plus ASCII compatibility and compactness, is why UTF-8 became the default of the web, of Linux, of JSON, of basically everything new.

> Default to UTF-8 everywhere and say so out loud. Set it on your files, your database columns, your HTTP `Content-Type` headers, your editor. The most reliable way to avoid Phase 1's mojibake is to make sure encoder and decoder both assume UTF-8 - and the way to guarantee that is to declare it explicitly rather than hope.

For builders: when you read or write text in code, name the encoding. `open(path, encoding="utf-8")`, not bare `open(path)` - the bare form uses the operating system's default, which differs between Windows and Linux and is a classic "works on my machine" trap. Make the byte format an explicit decision, not an accident of the host. (If you want the deeper picture of how those bytes sit in memory, the [/guides/how-a-computer-works](/guides/how-a-computer-works) guide covers the layer underneath.)

```quiz
[
  {
    "q": "What is the difference between Unicode and UTF-8?",
    "choices": ["They are two names for the same thing", "Unicode assigns a number (code point) to each character; UTF-8 is one way to encode those numbers as bytes", "Unicode is for English, UTF-8 is for everything else", "UTF-8 is older than Unicode"],
    "answer": 1,
    "explain": "Unicode is the catalogue mapping characters to code points. UTF-8 is a byte encoding - one of several ways to serialize those code points into bytes."
  },
  {
    "q": "How many bytes does the string \"café\" take when encoded as UTF-8?",
    "choices": ["4 bytes", "5 bytes - the é takes two", "8 bytes", "3 bytes"],
    "answer": 1,
    "explain": "Four characters, but the é is a 2-byte sequence (195, 169), so the total is 5 bytes. Byte count and character count are not the same in UTF-8."
  },
  {
    "q": "Why does UTF-8 have no byte-order (endianness) problem when UTF-16 does?",
    "choices": ["UTF-8 only stores English", "UTF-8's byte order is fixed by the format, while UTF-16 stores multi-byte units that can be ordered big- or little-endian", "UTF-8 always uses one byte", "UTF-16 is not a real encoding"],
    "answer": 1,
    "explain": "UTF-16 stores multi-byte code units whose order can differ (big- vs little-endian). UTF-8's byte sequence order is defined by the encoding itself, so the question never arises."
  }
]
```


---

# When Text Lies: Emoji, Graphemes, and the BOM

By now you have two solid ideas: bytes are not characters (Phase 1), and UTF-8 maps Unicode code points to a variable number of bytes (Phase 2). You might think the picture is complete: *byte → code point → character, done.* It isn't. There's one more layer, and it's the one that makes `len("👨‍👩‍👧")` return a number that looks insane. What a human calls "one character" and what Unicode calls "one code point" are **not the same thing**. This phase is about that gap, and about the small landmine called the byte-order mark.

## Three different lengths for "one" string

Take what looks like a single emoji: 👍🏽 (thumbs-up with a medium skin tone). Ask three different questions about its length and you get three different answers.

```text
👍🏽
  bytes (UTF-8):     8
  code points:       2   (base thumbs-up + a skin-tone modifier)
  graphemes:         1   (what the user sees and calls "one character")
```

*What just happened:* The same visible symbol is 8 bytes, 2 code points, or 1 grapheme depending on which layer you measure. The skin tone is a separate code point that *combines* with the base thumbs-up. A **grapheme** (or "grapheme cluster") is the real "character" a human perceives - and it can be built from several code points stuck together.

This is not an emoji curiosity. It is fundamental to how Unicode handles accents, scripts, and combining marks. The letter `é` can be stored two completely different ways:

```text
é  as one code point:   U+00E9                  (precomposed)
é  as two code points:  U+0065 (e) + U+0301 (◌́) (e + combining acute accent)
```

*What just happened:* Both render as an identical `é` on screen, but one is a single code point and the other is two code points glued together. They look the same and read the same to a human, yet a naive `==` comparison can say they are different strings, because their code points (and bytes) differ. This is why "the search box won't match the name that's obviously right there" bugs happen.

## Why string length lies

Most programming languages report string length in code points or, worse, in their internal code units - not in graphemes. So the number your code reports and the number a user would count diverge the moment emoji or combining marks appear.

```python runnable
s = "👍🏽"
print("len() reports:", len(s))
print("bytes:", len(s.encode("utf-8")))
# a human counts this as 1 character
```

*What just happened:* Python's `len` reports **2** for what the user sees as a single thumbs-up, because Python counts code points and this emoji is two of them (base + skin-tone modifier). It is 8 bytes on disk. None of these numbers is "1", which is the only answer a human would give. The lesson: **never use string length to count what a user sees as characters.** For tweets, SMS limits, password rules, or cursor movement, you must count graphemes, which usually means a dedicated library, not the built-in `len`.

> "How long is this string?" is an ambiguous question. Always answer the real one: bytes (for storage and network), code points (for Unicode processing), or graphemes (for anything a human reads). Picking the wrong one silently is how you ship a bug.

This three-layer model - bytes underneath, code points in the middle, graphemes on top - is the complete picture. If you have read the [/guides/data-structures-explained](/guides/data-structures-explained) guide, it is the same lesson as arrays versus the things stored in them: the container's count and the meaningful count are different questions.

## The byte-order mark: an invisible saboteur

There is one more gremlin that produces "impossible" bugs: the **byte-order mark**, or BOM. It is an optional invisible marker (the code point U+FEFF) that some tools - Windows Notepad and Excel are the usual culprits - stick at the very front of a UTF-8 file. In UTF-8 it is the three bytes `[239, 187, 191]`.

```text
file saved with BOM:
[239, 187, 191, 123, 34, ...]
 └──── BOM ────┘ └─ your actual "{"...

Naive reader sees the first character as:  "{   (with an invisible  before it)
```

*What just happened:* The reader, not knowing about the BOM, treats those three leading bytes as part of the content. Your JSON parser chokes because the file does not "start with `{`" - it starts with an invisible character. Your config key `name` is silently stored as `name`. Two files that look byte-for-byte identical in your editor behave differently because one has three ghost bytes at the front.

UTF-8 does not need a BOM (recall from Phase 2 that UTF-8 has no endianness to mark), so the cleanest rule is: **write UTF-8 without a BOM, and strip a BOM when reading if one sneaks in.** When a file mysteriously fails to parse on the very first character - especially a file that round-tripped through Excel or Notepad - check the first three bytes before you suspect anything else.

For builders: your debugging toolkit for any "weird text" bug is now three questions, asked in order. (1) *Wrong decoder?* - if whole runs of text are garbled, it is a Phase 1 mojibake mismatch; find who chose the encoding. (2) *Length surprise?* - if counts or comparisons are off near accents or emoji, you are confusing bytes, code points, and graphemes; pick the layer you actually mean. (3) *Fails on the first byte?* - suspect a BOM. Those three cover the overwhelming majority of text bugs you will ever meet.

```quiz
[
  {
    "q": "A user sees 👍🏽 as one character. Why does len() in many languages report 2?",
    "choices": ["The function is buggy", "It counts code points, and this emoji is a base character plus a separate skin-tone modifier (2 code points)", "It counts bytes", "Emoji always count as 2 for billing"],
    "answer": 1,
    "explain": "Length functions usually count code points (or internal code units), not graphemes. This emoji is 2 code points combined into 1 grapheme - the single symbol the user perceives."
  },
  {
    "q": "Two strings both display as \"é\" but a == comparison says they are different. What is the most likely cause?",
    "choices": ["A corrupted file", "One is the precomposed code point U+00E9 and the other is e + combining accent (two code points)", "Different fonts", "One is UTF-8 and one is ASCII"],
    "answer": 1,
    "explain": "The same visible é can be one precomposed code point or a base letter plus a combining mark. They render identically but have different code points and bytes, so a naive == fails."
  },
  {
    "q": "A JSON file fails to parse, complaining about the very first character, even though it clearly starts with {. What should you suspect first?",
    "choices": ["The JSON is invalid", "A byte-order mark (BOM) - three invisible bytes prepended by an editor like Notepad or Excel", "The disk is full", "JSON does not support objects"],
    "answer": 1,
    "explain": "A BOM (U+FEFF, the bytes 239 187 191 in UTF-8) is invisible but sits before your {. A naive parser treats it as content and fails on the first character. Write UTF-8 without a BOM."
  }
]
```
