# How a Neural Network Is Structured

> The structural anatomy of a neural network - layers, weights, biases, activation functions, and how one prediction flows through it - without touching how training works.


---

# How a Neural Network Is Structured

You've seen the diagram a hundred times: circles connected by lines, arranged in columns, labeled "input," "hidden," "output." But a diagram of dots and arrows doesn't tell you what's actually sitting inside each dot, or what happens to a number as it travels along one of those lines. This guide is entirely about that structure - what a neural network is built from, piece by piece, and how a single prediction moves through it start to finish. It deliberately stops short of *training* - how those connections get tuned is its own guide, and learning both at once is how this topic turns into mush. Here, we're just opening the case and looking at the parts.

## The phases

1. [Neurons, layers, and what "network" means](01-neurons-and-layers.md) - the layered picture: input, hidden, and output layers, and what one neuron structurally is.
2. [Weights, biases, and activation functions](02-weights-and-activations.md) - what a neuron actually computes, and why non-linearity is non-negotiable.
3. [The forward pass](03-the-forward-pass.md) - how one prediction flows through the whole structure, end to end.


---

# Neurons, Layers, and What "Network" Means

Strip away the buzzwords and a neural network is a structure for turning numbers into other numbers, arranged in a specific shape: numbers move through a sequence of stages, each doing a small transformation, until you're left with an answer. That's the whole shape. Everything else in this guide is detail on top of that one idea.

## The three kinds of layers

A neural network is organized into **layers** - groups of units arranged side by side, stacked one after another. Every network has exactly two layers with fixed jobs, and any number of layers in between:

- **The input layer** receives the raw data. If you're classifying an image, this is where the pixel values enter. If you're predicting house prices, this is where square footage, bedroom count, and location enter, each as a number. The input layer doesn't compute anything - it's just the network's on-ramp, one slot per piece of input data.
- **The output layer** produces the final answer. For a network that decides "cat or dog," this is two slots, each ending up holding something like a confidence score. For a network predicting a single price, it's one slot holding one number.
- **Hidden layers** sit between the two, and this is where the actual computation happens. They're called "hidden" because you never look at their values directly - you feed the input in one end and read the output at the other end, and what happens in between is intermediate work the network does for itself.

```mermaid
flowchart LR
  subgraph Input Layer
    I1((•))
    I2((•))
    I3((•))
  end
  subgraph Hidden Layer
    H1((•))
    H2((•))
    H3((•))
    H4((•))
  end
  subgraph Output Layer
    O1((•))
  end
  I1 --> H1
  I1 --> H2
  I2 --> H3
  I3 --> H4
  H1 --> O1
  H2 --> O1
  H3 --> O1
  H4 --> O1
```

*What this diagram means:* data enters on the left, one slot per input value, passes through a hidden layer where every unit is connected to every unit in the layer before it, and lands on the right as one or more output values. A network with more than one hidden layer stacked between input and output is what "deep learning" literally refers to - "deep" means "has multiple hidden layers," nothing more mystical than that.

📝 **Terminology:** the pattern where every unit in one layer connects to every unit in the next is called a **fully connected** (or **dense**) layer. It's the simplest and most common arrangement, though not the only one - networks built for images or sequences often use more specialized connection patterns. This guide sticks to the fully connected case, since it's the shape the rest of the anatomy is easiest to see in.

## What a single neuron structurally is

Zoom into one of those circles - a single **neuron** (also called a **unit** or **node**, all the same thing) - and structurally, it's just two things wired together:

1. A set of **inputs**, each arriving along one of the connecting lines from the previous layer.
2. A single **output**, sent forward along its own connecting lines to the next layer.

That's the entire structural shape: several numbers come in, one number goes out. Every neuron in a hidden or output layer works this way. What determines what number comes out - given the numbers that came in - is what Phase 2 is about. For now, the important thing is the shape: a neuron is a small unit with many inputs and one output, and a layer is a row of these units all doing their own version of the same job in parallel.

```text
Input layer   -> one slot per feature in your raw data, no computation
Hidden layer  -> the actual computation; can be one layer or many stacked
Output layer  -> one slot per number the network is meant to produce
```

## Why "network" and not just "layers"

The word "network" is doing real work here, not just sounding technical. A neuron doesn't only send its output to one place - it typically sends the exact same output value along a connection to *every* neuron in the next layer (in the fully connected case above). So the overall structure isn't a straight line of single connections; it's a dense mesh of connections between every pair of adjacent layers, which is exactly what a network - in the graph-theory sense, nodes connected by many edges - actually is.

This matters structurally because it means a single input value can influence *every* neuron in the next layer, and by extension, every neuron after that, all the way to the output. No one hidden neuron sees "the whole picture" of the input on its own, but by the time you reach the output layer, every output number has been shaped, in some tiny way, by every input number. That's the structural reason a network can represent something as complicated as "is this a picture of a cat" from raw pixel values: not because any one neuron is smart, but because there are enormous numbers of these small connections layered on top of each other.

The remaining question - and it's the one that actually gives a neuron its computational power - is what exactly a neuron *does* with the numbers arriving on all those input connections before it produces its one output number. That's Phase 2.


---

# Weights, Biases, and Activation Functions

Phase 1 left one question hanging: given several numbers arriving at a neuron, how does it turn them into the one number it sends onward? The answer has exactly two steps, done in order, every time, by every neuron in every hidden and output layer. Once you have those two steps, you have the entire computational engine of a neural network - the rest is just this pattern, repeated at scale.

## Step one: a weighted sum

Every connection arriving at a neuron carries a **weight** - a number that says how much importance that particular input gets. A neuron doesn't just add up its inputs plainly; it multiplies each input by its own weight first, then adds up the results.

```text
weighted sum = (input_1 × weight_1) + (input_2 × weight_2) + (input_3 × weight_3) + ...
```

*What this means:* if `weight_1` is large, that input has a big say in the neuron's behavior. If a weight is near zero, that input is effectively being ignored, no matter what value it carries. If a weight is negative, that input actually pushes the sum *down* rather than up. Structurally, the weights are what encode everything the network has "learned" - how those specific numbers get set is the subject of a separate guide, but the two things to hold onto here are: every connection has its own independent weight, and the weighted sum is just each input scaled by how much it matters, added together.

On top of the weighted sum, a neuron adds one more number: the **bias**. It's a single value, one per neuron, added after the weighted sum:

```text
neuron's raw value = weighted sum + bias
```

The bias exists to shift the neuron's output up or down independent of its inputs - think of it as the neuron's baseline tendency, the value it produces even when every input is zero. Without a bias, every neuron would be stuck passing exactly through zero when its inputs are all zero, and there'd be no way to represent something like "this neuron should activate readily" versus "this neuron should be hard to trigger" as a separate, tunable knob from the weights themselves.

## Step two: squash it through an activation function

Here's where it gets interesting. If a neuron stopped at the weighted sum plus bias, its output would just be one plain number, capable of growing arbitrarily large or arbitrarily negative depending on the inputs. Instead, that raw value gets passed through an **activation function** - a fixed mathematical shape that takes the raw number in and produces the neuron's actual output.

A few common activation functions, structurally:

- **ReLU** (Rectified Linear Unit) - the simplest and most widely used today. It passes positive values through unchanged, and turns any negative value into zero. Structurally: "if it's positive, keep it; if it's negative, kill it."
- **Sigmoid** - squashes any input, however large or small, into a value between 0 and 1. Useful when you want a neuron's output to look like a probability.
- **Tanh** - similar to sigmoid in shape, but squashes into a range between -1 and 1 instead of 0 and 1.

```mermaid
flowchart LR
  A["Inputs × weights, + bias"] --> B["Raw value<br/>(could be any number)"]
  B --> C["Activation function"]
  C --> D["Neuron's actual output<br/>(bounded / reshaped)"]
```

*What this diagram means:* the weighted sum and bias produce one unshaped number, and the activation function is a deliberate, fixed reshaping step applied to that number before it's allowed to leave the neuron. Different activation functions reshape it differently, but every neuron in a typical hidden layer applies the same one, consistently, every time.

## Why the activation function isn't optional

This is the part that looks like a minor implementation detail and is actually load-bearing for the entire idea of a "deep" network. Here's the problem if you skip it: a weighted sum plus a bias, with nothing else applied, is a **linear** operation - scaling and adding, nothing more. And here's the fact about linear operations that breaks everything: stacking linear operations on top of each other, no matter how many layers you use, produces something that is *still just one linear operation* overall.

Concretely: if every neuron in every layer only ever computed a weighted sum with no activation function, a 50-layer network would behave *exactly* the same as some single, one-layer network doing one weighted sum - mathematically collapsible into it, every time. All that depth, all those neurons, all those weights, would buy you nothing beyond what a single layer could already do. You could stack a thousand purely linear layers and it would still only be capable of representing straight lines and flat planes - never a curve, never a genuine "if this AND that, but not the other thing" kind of decision boundary.

> Stacking linear layers without a non-linear activation function between them is mathematically pointless - no matter how many you stack, the result collapses into one single linear function.

The **non-linear** activation function is what breaks that collapse. Because ReLU, sigmoid, and similar functions bend the numbers - they don't just scale and shift, they reshape the relationship - stacking layers *with* a non-linearity between them genuinely builds something new at each layer, something the previous layers alone couldn't already express. This is the actual reason "deep" networks with many layers can represent extremely complicated relationships (recognizing a face, understanding language) that a single layer fundamentally cannot: not because there are more numbers involved, but because each non-linear bend lets the next layer build on a genuinely more complex shape than the one before it.

```quiz
[
  {
    "q": "What are the two things a neuron computes, in order, before producing its output?",
    "choices": [
      "A random number, then a fixed lookup",
      "A weighted sum plus a bias, then the result passed through an activation function",
      "An average of its inputs, then a rounding step",
      "A comparison against every other neuron in the layer"
    ],
    "answer": 1,
    "explain": "Every neuron multiplies each input by its weight, sums the results, adds its bias, and then reshapes that raw value through an activation function."
  },
  {
    "q": "What does a weight control in a neural network?",
    "choices": [
      "How many layers the network has",
      "How much a specific input contributes to a neuron's weighted sum",
      "The order in which neurons are computed",
      "Whether a neuron is part of the input or output layer"
    ],
    "answer": 1,
    "explain": "Each connection has its own weight, which scales that specific input's contribution - a near-zero weight effectively ignores the input, a negative weight pushes the sum down."
  },
  {
    "q": "What happens if you remove the non-linear activation function from every neuron in a multi-layer network?",
    "choices": [
      "Nothing changes; activation functions are purely cosmetic",
      "The network trains faster with no downside",
      "The entire stack of layers collapses into a single equivalent linear function, no matter how many layers there are",
      "The network can only produce negative numbers"
    ],
    "answer": 2,
    "explain": "Stacked linear operations (weighted sums with no non-linearity) always reduce to one linear operation overall. The non-linear activation function is what lets depth actually add representational power."
  }
]
```

Watch it animated: [how a neural network is structured](/explainers/NeuralNetwork.dc.html)


---

# The Forward Pass

You now have every piece: layers arranged input to output, neurons that each compute a weighted sum plus a bias and reshape it through an activation function, and connections carrying values from one layer to the next. This phase puts it all in motion - following one prediction as it travels through the whole structure, start to finish.

## What "forward pass" means

The **forward pass** is the name for this entire journey: feeding one set of input values into the input layer, and following the computation, layer by layer, until a final value emerges from the output layer. It's called "forward" because data only ever moves in one direction through it - from input toward output, never backward - during this particular operation. (Tuning the weights, covered elsewhere, does involve information flowing backward through the network - which is exactly why that direction gets its own name and its own guide, rather than being folded into this one.)

## Walking through one pass, layer by layer

Say you have a small network: 3 input values, one hidden layer with 4 neurons, and 1 output neuron - structurally the same shape from Phase 1's diagram. Here's what actually happens to a single set of inputs:

```text
Step 1: The 3 input values enter the input layer unchanged.
         (e.g. square footage, bedroom count, distance to downtown)

Step 2: Each of the 4 hidden neurons receives all 3 input values.
         Each hidden neuron independently computes:
           its own weighted sum of the 3 inputs, using its own weights
           + its own bias
           -> passed through its activation function
         Result: 4 output numbers, one per hidden neuron.

Step 3: The single output neuron receives all 4 hidden-layer outputs.
         It computes:
           its own weighted sum of those 4 values, using its own weights
           + its own bias
           -> passed through its activation function (or none, for some outputs)
         Result: 1 final number - the network's prediction.
```

*What just happened:* the same two-step recipe from Phase 2 - weighted sum plus bias, then activation function - ran once per neuron, and it ran in a strict order: every neuron in the hidden layer had to finish before the output neuron could start, because the output neuron's inputs *are* the hidden layer's outputs. Nothing skips ahead. Each layer fully finishes its computation before the next layer can begin, which is exactly what makes this a well-defined, one-directional flow rather than a tangle.

```mermaid
sequenceDiagram
  participant Input as Input Layer
  participant Hidden as Hidden Layer
  participant Output as Output Layer
  Input->>Hidden: 3 raw values
  Hidden->>Hidden: Each neuron: weighted sum + bias -> activation
  Hidden->>Output: 4 computed values
  Output->>Output: Weighted sum + bias -> activation
  Output-->>Input: 1 final prediction
```

*What this diagram means:* the whole forward pass is a strict relay - each layer waits for the complete output of the layer before it, transforms it, and hands off a new set of values to the next layer. By the time the process reaches the arrow back to the start, that's purely showing "this is the answer that came out," not data actually flowing backward.

## Why this is called "inference"

When a trained network is used to make a prediction on new data - a photo it's never seen, a house it's never priced - running the forward pass on that input is called **inference**. It's worth being precise about what is and isn't happening during inference: the weights and biases throughout the network are already fixed numbers at this point, set once beforehand. A forward pass, including every inference request your app ever makes to a trained model, does not change a single weight. It's purely a read-and-compute operation: take the fixed weights, take the new input, run the relay described above, get an answer.

> A forward pass never changes the network. It's the network, exactly as it currently is, computing one answer for one input.

This is also why the same input, run twice through the same trained network, always produces the exact same output - there's no randomness or memory involved in the structure itself (some networks deliberately introduce randomness during training as a technique, but the raw forward-pass mechanism described here is a fixed, repeatable computation).

## Where the weights actually come from

Everything in this guide assumed the weights and biases already had sensible values - Phase 2 talked about what a weight *does*, never about how it ends up being 0.73 instead of some other number. That's a deliberate boundary. How a network starts with random, useless weights and gradually adjusts them until the forward pass described above actually produces good predictions is a genuinely separate topic, involving comparing the network's output to the right answer and pushing the weights in a direction that reduces the error. If you want that half of the picture, that's exactly what [how a model learns](/guides/how-a-model-learns) covers - how the weights actually get set is a separate topic from the structure that carries them.

For now, the anatomy is complete: layers of neurons, each computing a weighted sum plus a bias and reshaping it through a non-linear activation function, relaying values forward from input to output. That structure is the same whether the network is freshly initialized with random junk or fully trained and state-of-the-art - training only ever changes the numbers sitting inside this same shape.
