# System Calls Explained

> The controlled doorway between your program and the kernel - why it exists, what happens during one, and why it matters for real performance.


---

# System Calls Explained

Every time your program reads a file, sends a network packet, or even asks what time it is, it can't do that work itself. It has to ask the kernel - the core of the operating system - through a tightly controlled doorway called a **system call**. That doorway exists for a reason, costs something every time you walk through it, and understanding both explains a surprising number of real performance decisions you'll run into as a developer.

## How to read this

Read it in order. Phase 1 explains why programs can't touch hardware directly. Phase 2 walks through the mechanics of a syscall: the trap, the mode switch, the return. Phase 3 gets practical - why the cost of a mode switch is real, why buffering exists, and how tools like `strace` let you watch syscalls happen live. The mechanism is the same shape on Linux, macOS, and Windows, even though the exact instructions and tool names differ.

## The phases

1. [Why programs can't touch hardware directly](01-user-mode-vs-kernel-mode.md) - the wall between user mode and kernel mode, and why it exists.
2. [What actually happens during a syscall](02-the-mechanics-of-a-syscall.md) - the trap, the mode switch, the syscall number and arguments, the return.
3. [Why syscalls matter for real performance](03-why-syscalls-matter-for-performance.md) - the cost of a mode switch, why you buffer, and watching syscalls with strace.


---

# Why programs can't touch hardware directly

When you write `print("hello")` or open a file, it feels like your program does it directly. But nothing in your program actually touches the disk, the network card, or the screen. A wall sits between your code and the hardware, enforced by the CPU itself, and your program has to ask permission to cross it every single time. That wall has a name: the boundary between **user mode** and **kernel mode**.

## Two modes, enforced by the CPU

Modern CPUs support at least two distinct privilege levels, enforced in hardware, not just software convention. **Kernel mode** (sometimes called supervisor mode or ring 0) can execute any instruction, including the dangerous ones: directly manipulating hardware registers, managing memory for every process on the machine, halting the CPU. **User mode** (ring 3, on x86) is deliberately restricted - a whole category of instructions refuse to execute there at all, and the CPU raises a fault if you try.

```text
Kernel mode:  can do anything - talk to hardware, manage all memory, control other processes
User mode:    restricted - normal arithmetic and logic only, no direct hardware access
```

*What just happened:* every regular program you run - your browser, your text editor, your own code - runs in user mode. Only the kernel itself runs in kernel mode. This isn't a polite convention your program agrees to follow; it's physically enforced by the processor. Try to execute a privileged instruction from user mode and the CPU refuses, generating a hardware fault instead of running it.

## Why the wall exists at all

Picture a machine with no such wall: any program could directly write to any part of memory, including memory belonging to other programs, or issue raw commands to the disk controller, or halt the entire CPU. One careless or malicious program could crash the whole system or read another program's private data with nothing standing in the way.

```text
No wall:    Program A can read/write Program B's memory directly. One bug crashes everything.
With wall:  Program A can only touch its own memory. Anything shared must go through the kernel.
```

*What just happened:* this is **protection**, and it's the entire reason the wall exists. The kernel is the one piece of software trusted to touch shared, sensitive resources - physical memory, disks, network hardware - precisely because it's the one piece of software that can enforce rules about *who* gets to touch *what*. Take that arbitration away and any bug anywhere can corrupt anything.

> The wall isn't there to slow you down. It's there so that a bug in your program stays a bug in your program, instead of becoming a bug in every program running on the machine.

## So how does a program get anything done?

If user mode can't touch hardware directly, your program needs a way to ask the kernel - which *can* - to do something on its behalf. That controlled crossing point is the **system call**: a program in user mode makes a specific, well-defined request, the CPU switches to kernel mode for exactly as long as the kernel needs to handle it, and control returns to user mode with the result.

```text
Your program (user mode): "please read 100 bytes from this file"
       |
       v  (system call - the only sanctioned way across the wall)
Kernel (kernel mode): actually talks to the disk controller, reads the bytes
       |
       v
Your program (user mode): receives the 100 bytes, keeps running
```

*What just happened:* your program never touches the disk controller. It asks, through the one narrow, audited doorway the kernel exposes, and the kernel does the actual hardware work on its behalf. Every file read, every network send, every memory allocation beyond what your program already owns, every `time()` call - all of it goes through this same doorway.

## The mental model to keep

Think of user mode and kernel mode as two rooms with exactly one door between them, controlled by the kernel. Your program lives entirely in the user-mode room and can do plenty on its own there - arithmetic, string manipulation, working with data already in its own memory. The moment it needs something from outside that room - hardware, another process, memory it doesn't yet own - it has to knock on that door and wait for the kernel to open it, do the work, and close it again.

That knock has a name and a very specific mechanism, and that's exactly what Phase 2 walks through: what actually happens, instruction by instruction, in the moment a system call fires.


---

# What actually happens during a syscall

Phase 1 established that a system call is the only sanctioned way across the user-mode/kernel-mode wall. This phase opens that mechanism up: the specific steps between your program calling `read()` and getting bytes back.

## Step 1: your program requests a specific syscall, by number

Every system call the kernel supports has a fixed, agreed-upon number. `read` might be syscall number 0, `write` might be 1, `open` might be 2 - the exact numbers differ by operating system, but the idea is universal: there's a numbered table, and your program's C library (or runtime) knows which number corresponds to which operation.

```text
read(fd, buffer, count)
  -> your language's standard library translates this into:
     syscall number: 0 (on Linux x86-64, for example)
     arguments: fd, buffer address, count
```

*What just happened:* when you call a high-level function like `read()`, a thin layer underneath - the C standard library, or your language runtime - packages that call into the specific syscall number and arguments the kernel expects. You almost never deal with the raw number yourself; the library handles the translation.

## Step 2: the trap - a deliberate software interrupt

Here's the part that makes the mode switch possible. Your program executes a special CPU instruction whose entire purpose is to voluntarily trigger a switch into kernel mode - on x86-64 Linux this is the `syscall` instruction (older systems used a software interrupt, `int 0x80`). This is called a **trap**: not an error, but a deliberate signal that says "kernel, take over from here."

```text
user mode:   ... normal instructions ...
             syscall            <- special instruction: trap into kernel mode
kernel mode: (CPU privilege level switches, kernel's trap handler runs)
```

*What just happened:* this single instruction is the entire crossing point. Before it executes, the CPU is in user mode and restricted. The instant it executes, the CPU privilege level flips to kernel mode, and control jumps to a fixed, known entry point inside the kernel - the kernel's trap handler, waiting specifically for this. Nothing about this is negotiable from the user-mode side; the program can't pick where in the kernel execution lands, which is itself part of the protection Phase 1 described.

## Step 3: the kernel reads the syscall number and dispatches

Once inside the trap handler, the kernel looks at the syscall number your program placed in a specific CPU register before trapping, and uses it to look up which internal kernel function handles that request - a **syscall table**, essentially an array of function pointers indexed by syscall number.

```text
syscall_table[0]  -> sys_read()
syscall_table[1]  -> sys_write()
syscall_table[2]  -> sys_open()
...

kernel: number = 0 -> calls sys_read(fd, buffer, count)
```

*What just happened:* the kernel doesn't guess or parse anything fuzzy - it does a direct lookup by number and calls the corresponding function, passing along the arguments your program provided. `sys_read()` is real kernel code that knows how to talk to the filesystem layer and, eventually, the actual disk driver, entirely in kernel mode where it's allowed to do so.

## Step 4: the kernel does the actual work

Now the requested work actually happens - with full kernel privileges. For `read()`, that means checking the file descriptor is valid, checking permissions, asking the filesystem where those bytes live on disk, and asking the disk driver to fetch them. None of this could have happened in user mode; this is precisely the privileged work Phase 1 said only the kernel can do.

## Step 5: the return - switching back to user mode

Once the kernel finishes the work (or fails - the same path handles both cases, with a return value indicating error or success), it executes a return-from-trap instruction. This flips the CPU back to user mode and resumes your program exactly where it left off, right after the `syscall` instruction, with the result available.

```text
kernel mode: sys_read() finishes, places result in a register
             sysret / iret         <- return-from-trap: flips back to user mode
user mode:   ... program resumes here, with the read bytes now available ...
```

*What just happened:* your program's execution wasn't destroyed or restarted - it was paused at one precise instruction boundary, the kernel did privileged work on its behalf, and execution resumed at the next instruction as if nothing unusual happened, except now a buffer that was empty is full of file data.

> A syscall is not a function call to some code living in your own process. It's a full round trip out of your program's privilege level, into the kernel, and back - with a hardware-enforced wall crossed twice, once in each direction.

## The whole sequence, together

```text
1. Your code calls a library function (e.g. read())
2. Library sets up the syscall number + arguments in specific registers
3. `syscall` instruction traps into kernel mode
4. Kernel looks up the syscall number in its table, dispatches to the handler
5. Handler does the privileged work (touches the filesystem/disk/network)
6. Kernel returns; CPU switches back to user mode
7. Your program resumes with the result
```

*What just happened:* every one of your program's interactions with the outside world - every file, every socket, every millisecond of wall-clock time it asks for - runs this exact seven-step sequence. It happens so often, and usually so fast, that you never see it directly. But it isn't free, and that cost is the subject of Phase 3.


---

# Why syscalls matter for real performance

Phase 2 walked through the mechanics: trap, mode switch, dispatch, work, return. Each step takes real time - not a huge amount on any single call, but enough to shape how experienced developers write I/O code. This phase covers why that small cost adds up, and how to actually watch it happening on a running program.

## The cost of a single mode switch

A system call is slower than a regular function call in your own code, and the reason traces directly back to Phase 2's mechanics. A normal function call is just a jump to a nearby address, still in user mode, with the CPU's instruction pipeline and caches barely disturbed. A syscall does considerably more: it saves your program's register state, flips the CPU's privilege level, jumps into kernel code (which likely isn't sitting in the same CPU cache lines your program was just using - a **cache-cold** jump), does its work, and reverses the whole thing on the way back.

```text
Regular function call:   nanoseconds - stays in user mode, cache-friendly
System call:              tens to low hundreds of nanoseconds - mode switch + kernel dispatch overhead,
                          even before the kernel does any actual work
```

*What just happened:* even a syscall that does almost nothing - like asking for the current time - pays this fixed overhead just for crossing the wall and coming back. The kernel isn't slow at its job; crossing the boundary at all has an unavoidable fixed cost, separate from whatever work happens once you're across.

## Why you buffer instead of calling write() once per byte

This is the most common place this cost shows up in real code. Imagine writing a file one byte at a time:

```text
# the slow way - one syscall per byte
for byte in data:
    write(fd, byte, 1)      # each call: full trap, mode switch, kernel work, return
```

*What just happened:* if `data` is a megabyte, that's roughly a million system calls, each paying the full mode-switch overhead from the section above - on top of whatever the actual disk write costs. The fixed per-call overhead, multiplied a million times, can dwarf the cost of the real work being done.

Compare that to buffering: accumulate data in memory (cheap, cache-friendly, no mode switch) and issue one `write()` call for a large chunk at once.

```text
# the fast way - accumulate in a user-mode buffer, one syscall for the whole chunk
buffer = []
for byte in data:
    buffer.append(byte)     # plain memory write, no syscall, no mode switch
write(fd, buffer, len(buffer))   # one syscall for the entire megabyte
```

*What just happened:* you paid the fixed mode-switch cost exactly once instead of a million times. This is why every serious I/O library - file handles, network sockets, standard output - buffers by default. `printf` doesn't call `write()` for every character you print; it accumulates output and flushes in larger chunks. The N+1-style trap here has the same shape as issuing one database query per row instead of one query for all of them: the fix in both cases is batching the expensive boundary crossing instead of paying its fixed cost over and over.

> The lesson isn't "syscalls are bad." It's that crossing the user/kernel wall has a real, fixed cost per crossing - so the number of crossings matters as much as the total amount of work being done across them.

## Seeing syscalls actually happen: strace and friends

Because syscalls are the exact point where your program touches the outside world, tracing them is one of the most reliable ways to understand what a program is *actually* doing, independent of what its source code claims. On Linux, the tool is `strace`; macOS has `dtruss` (built on DTrace); Windows has Process Monitor for similar visibility into system-level activity.

```text
$ strace -c ./my_program

% time     seconds  usecs/call     calls    syscall
------ ----------- ----------- --------- ----------------
 45.20    0.003821          12       318 write
 30.11    0.002544           8       318 read
 12.03    0.001017          32        32 openat
  ...
```

*What just happened:* `strace -c` runs the program and prints a summary of every syscall it made, how many times, and how much time was spent inside the kernel handling each one. Seeing `write` called 318 times when you expected 3 is often the exact moment you discover an unbuffered loop like the one earlier in this phase - the syscall count is direct, unambiguous evidence, not a guess.

You can also trace without the summary, seeing each call as it happens with its actual arguments:

```text
$ strace ./my_program
openat(AT_FDCWD, "config.json", O_RDONLY) = 3
read(3, "{\"key\": \"value\"}", 4096) = 17
close(3)                               = 0
```

*What just happened:* this is the seven-step sequence from Phase 2, made visible from outside the process - every trap into the kernel, with its arguments and return value, printed in order. When a program is mysteriously slow, hanging, or touching files it shouldn't, this is often the fastest way to find out what it's really doing, because it bypasses whatever the source code claims and shows the actual boundary crossings as they happen.

## Bringing it together

The wall between user mode and kernel mode exists for protection. Crossing it is a specific, mechanical process - trap, mode switch, dispatch, work, return - and that process has a real, fixed cost independent of how much work happens on the other side. Once you internalize that cost, a lot of otherwise-mysterious performance advice stops being folklore: buffer your I/O, batch your writes, and when something's inexplicably slow, look at what it's actually asking the kernel to do.

Watch it animated: [system calls](/explainers/SystemCalls.dc.html)

```quiz
[
  {
    "q": "Why is a system call slower than a normal function call in your own code?",
    "choices": [
      "Syscalls always involve a disk, which is inherently slow",
      "It requires a CPU privilege mode switch and a jump into kernel code, on top of whatever work is actually done",
      "The kernel deliberately adds a delay for security reasons",
      "Syscalls are only slow on older CPUs"
    ],
    "answer": 1,
    "explain": "The mode switch itself - saving state, flipping privilege level, cache-cold jump into the kernel, and back - has a fixed cost separate from the actual work performed."
  },
  {
    "q": "Why do I/O libraries buffer output instead of calling write() for every byte?",
    "choices": [
      "Buffering makes the disk itself write faster",
      "It avoids paying the fixed mode-switch overhead once per byte by batching many bytes into one syscall",
      "The kernel rejects small writes",
      "Buffering is only needed for network sockets, not files"
    ],
    "answer": 1,
    "explain": "Each syscall pays the same fixed crossing cost regardless of how much data it carries - batching into fewer, larger calls amortizes that cost across far more data."
  },
  {
    "q": "What does a tool like strace actually show you?",
    "choices": [
      "The source code of the program being traced",
      "A prediction of future performance based on static analysis",
      "The real system calls the program makes as it runs, with arguments and return values",
      "Only the syscalls that fail"
    ],
    "answer": 2,
    "explain": "strace intercepts and prints the actual trap-into-kernel events as they happen, showing unambiguous evidence of what a program is really doing, independent of its source code."
  }
]
```
