# Calculus: The Math of How Things Change

> Calculus is not about memorizing derivative rules. It is the math of change: how fast something is moving, how to find the best possible outcome, how to add up a million tiny pieces. This guide starts from a car's speedometer and builds up to the ideas behind machine learning gradients and physics simulations.


---

# Calculus: The Math of How Things Change

If you have ever watched a car's speedometer, used a learning rate in a machine learning model, or seen a graph of "users over time" in a product dashboard, you have used calculus. The difference is that the computer knew it was using calculus, and you did not.

This guide fixes that. We are not going to drill the quotient rule until your eyes bleed. We are going to start from something you already understand - a speedometer - and build up to the ideas that make calculus useful. By the end, a derivative will look like "how fast, right now," and an integral will look like "the total so far." That is all they are.

This is the ninth guide in the Mathematics track. It assumes the function idea from [Sets, Relations, and Functions](/guides/sets-relations-and-functions) and the algebra of lines from [Numbers & Number Systems](/guides/numbers-and-number-systems). If you can read a graph and understand slope, you are ready.

## How to read this
- **Here for the "what is calculus actually for" answer?** Start with [Phase 1](01-derivatives-as-right-now-speed.md) - derivatives as instantaneous speed.
- **Want the full toolkit?** Read in order - optimization builds on derivatives, and integrals build on the idea of accumulation.

## The phases
1. **[Derivatives as Right Now Speed](01-derivatives-as-right-now-speed.md)** - slope, velocity, marginal cost, and the code that approximates a derivative from data points.
2. **[Optimization and What Is the Best I Can Do](02-optimization-and-whats-the-best-i-can-do.md)** - finding maxima and minima, gradient descent intuition, and how a neural network "learns" by nudging parameters.
3. **[Integrals as the Total So Far](03-integrals-as-the-total-so-far.md)** - area under a curve, expected value, Riemann sums, and how profiling data is an integral over time.

> This builds on [Sets, Relations, and Functions](/guides/sets-relations-and-functions) (functions as mappings) and [Counting & Combinatorics](/guides/counting-and-combinatorics) (summation as repeated addition). It is the continuous math behind much of modern computing.


---

# Derivatives as Right Now Speed

## The speedometer that taught me everything

You are driving down a highway. The speedometer reads 65 miles per hour. That number is a derivative.

It tells you how fast you are moving *right now*. Not how far you traveled in the last hour. Not how far you will travel in the next hour. Only this instant. That is what a derivative is: the rate of change at a single point in time.

If you press the gas, the speedometer climbs. If you let off, it falls. The speedometer is tracking the derivative of your position with respect to time.

## From average speed to instantaneous speed

Suppose you drive 100 miles in 2 hours. Your average speed is 50 miles per hour. But you probably were not going exactly 50 the whole time. You accelerated, cruised, maybe stopped for coffee.

Average speed is easy: total distance divided by total time. Instantaneous speed is harder. It is the speed at a single moment, and it requires a new idea.

Here is the trick. To find the speed at exactly 1 hour, look at a very short interval around 1 hour - say from 1 hour to 1.0001 hours. Compute the average speed over that tiny interval. As the interval gets smaller, the average speed gets closer and closer to the true instantaneous speed.

The derivative is the limit of that process as the interval shrinks to zero. It is the slope of the tangent line to the curve at that point.

## The slope of a hill

Picture a hill. At the bottom, the hill is flat - the slope is zero. As you climb, the slope gets steeper. At the top, the slope is zero again. On the way down, the slope is negative.

The slope at any point is the derivative of the height function at that point. If the height function is `h(x)`, then the derivative `h'(x)` tells you how steep the hill is at position `x`.

For a straight line, the slope is constant. For a curve, the slope changes from point to point. The derivative is the function that tells you the slope at every point.

## The derivative of a simple function

Consider the function `f(x) = x^2`. This is a parabola. At `x = 0`, the slope is zero (the bottom of the valley). At `x = 1`, the slope is 2. At `x = 2`, the slope is 4.

The derivative of `x^2` is `2x`. That is the rule: to find the slope at any point, double the x value.

```
f(x) = x^2
f'(x) = 2x
```

At `x = 3`, the slope is `2 * 3 = 6`. The curve is rising 6 units for every 1 unit you move to the right.

You do not need to memorize this rule. You need to understand what it means: the derivative of a function is a new function that gives the slope of the original function at every point.

## Marginal cost: the economics of one more

In economics, the **marginal cost** is the cost of producing one more unit. If you have made 100 widgets and you are considering the 101st, the marginal cost is the derivative of the total cost function at 100.

Suppose the total cost of producing `x` widgets is `C(x) = x^2 + 10x + 50`. The marginal cost at `x = 100` is `C'(100) = 2 * 100 + 10 = 210`, so the 101st widget costs about $210. (The exact cost of that one unit, `C(101) - C(100)`, is $211 - the derivative is the instantaneous rate, a fast and very close estimate.)

The derivative turns a question about totals ("how much does 100 widgets cost?") into a question about rates ("how much does one more cost?"). That translation is the whole point of calculus.

## Approximating a derivative from data

In the real world, you often have data points instead of a formula. You can approximate the derivative by computing the slope between two nearby points.

```
f'(x) ≈ (f(x + h) - f(x)) / h
```

The smaller `h` is, the better the approximation. This is called a **finite difference**, and it is how numerical libraries compute derivatives when they do not have an analytic formula.

## See it run

Here is a function and its derivative, computed both by the rule and by finite differences.

```python runnable
def f(x):
    return x ** 2

def derivative_at(x, h=0.0001):
    return (f(x + h) - f(x)) / h

for x in [0, 1, 2, 3, 4]:
    print(f"f({x}) = {f(x)}, f'({x}) ≈ {derivative_at(x):.4f}, exact = {2 * x}")
```

*What just happened:* The function `f(x) = x^2` was evaluated at several points. The `derivative_at` function approximated the derivative using a finite difference with `h = 0.0001`. The results match the exact derivative `2x` to four decimal places. The approximation gets better as `h` gets smaller.

## For builders

Derivatives are not only for math class. They are the engine behind much of modern computing.

- **Machine learning** - Training a neural network is minimizing a loss function. The gradient of the loss function tells you which direction to nudge the parameters to reduce the loss. That nudge is a derivative.
- **Physics simulations** - In a game or physics engine, the position of an object changes over time. The velocity is the derivative of position. The acceleration is the derivative of velocity. Simulating motion means integrating these derivatives over time.
- **Optimization** - Any time you want to find the maximum or minimum of a function - the best learning rate, the cheapest production level, the fastest route - you are looking for where the derivative is zero.
- **Signal processing** - Filters that detect edges in an image or changes in an audio signal are computing derivatives. An edge is a place where pixel intensity changes rapidly, which means the derivative is large.

> The key insight: a derivative is "how fast, right now." It turns a static picture into a movie, a total into a rate, a cost into a marginal cost. Once you see it that way, the formulas stop being magic and start being useful.

## What we have built

- A **derivative** is the rate of change of a function at a single point.
- For `f(x) = x^2`, the derivative is `f'(x) = 2x`.
- The derivative at a point is the slope of the tangent line to the curve at that point.
- **Marginal cost** is the derivative of the total cost function.
- A **finite difference** approximates a derivative from data points: `(f(x + h) - f(x)) / h`.
- In code, derivatives power machine learning, physics simulations, optimization, and signal processing.

A quick check before you move on:

```quiz
[
  {
    "q": "What does the derivative of a function represent?",
    "choices": ["The total change over a long interval", "The rate of change at a single point", "The average value of the function", "The area under the curve"],
    "answer": 1,
    "explain": "The derivative is the rate of change at a single point. For position with respect to time, it is instantaneous velocity. For cost with respect to quantity, it is marginal cost. It is not an average over an interval."
  },
  {
    "q": "If f(x) = x^2, what is f'(3)?",
    "choices": ["3", "6", "9", "12"],
    "answer": 1,
    "explain": "The derivative of x^2 is 2x. At x = 3, f'(3) = 2 * 3 = 6. The slope of the curve y = x^2 at the point (3, 9) is 6."
  },
  {
    "q": "How do you approximate a derivative from data points?",
    "choices": ["Divide the total change by the number of points", "Take the average of all the points", "Compute the slope between two nearby points: (f(x + h) - f(x)) / h", "Find the highest point and subtract the lowest"],
    "answer": 2,
    "explain": "A finite difference approximates the derivative by computing the slope between two points that are very close together. The smaller the distance h, the better the approximation of the true instantaneous rate of change."
  }
]
```


---

# Optimization and What Is the Best I Can Do

## The hill you are climbing

Picture yourself on a hill in the fog. You cannot see the top or the bottom. You can only feel the slope under your feet. Your goal is to reach the highest point.

What do you do? You look at the slope. If the ground rises to your right, you step right. If it rises to your left, you step left. You keep stepping in the direction that goes up, and eventually you reach a peak.

That is optimization. You have a function - the height of the hill - and you want to find its maximum. The derivative tells you which way is up. You follow it until you can go no higher.

The same idea works for finding minima. If you want the lowest point in a valley, you follow the slope downward.

## Where the derivative is zero

At the very top of a hill, the ground is flat. The slope is zero. At the very bottom of a valley, the ground is also flat. The slope is zero.

These flat spots are called **critical points**. They are the only places where a smooth function can have a maximum or a minimum. To find the best outcome, you find where the derivative is zero and then check whether that point is a peak, a valley, or a flat plain.

```
f(x) = x^2
f'(x) = 2x
Set f'(x) = 0: 2x = 0, so x = 0.
At x = 0, f(0) = 0. This is the minimum of the parabola.
```

The derivative told us exactly where the minimum is. No guessing, no plotting a hundred points. One equation, one solution.

## Local maxima vs global maxima

Not every flat spot is the best spot. A hill can have many local peaks. Each local peak is higher than the ground immediately around it, but a distant mountain may be taller.

In optimization, a **local maximum** is the best point in its neighborhood. A **global maximum** is the best point overall. Finding a local maximum is easy: follow the slope upward until it flattens. Finding the global maximum is hard, because you might have to explore the whole landscape.

This is the central challenge of machine learning. A neural network has millions of parameters. The loss function - the measure of how wrong the network is - is a surface in a million-dimensional space. Training the network means finding the lowest point on that surface. But the surface has many local valleys, and gradient descent can get stuck in one of them.

## Gradient descent: the algorithm that learns

**Gradient descent** is the workhorse of machine learning. It is the algorithm that trains neural networks, fits regression models, and optimizes recommendation systems.

The idea is identical to the hill-climbing metaphor, but in higher dimensions. Instead of a slope (one number), you have a gradient (a vector of partial derivatives). The gradient points in the direction of steepest ascent. To minimize a function, you step in the opposite direction.

```
new_position = old_position - learning_rate * gradient
```

The **learning rate** controls how big a step you take. Too small, and you crawl. Too large, and you overshoot the minimum, bouncing back and forth or even flying off to infinity.

Each step reduces the loss a little. Repeat thousands of times, and the model converges to a local minimum. That is "learning."

## A tiny example: fitting a line

Suppose you have data points that roughly follow a line. You want to find the slope and intercept that make the line fit the data best. The "best fit" is the line that minimizes the sum of squared errors - the distance between each data point and the line.

This is an optimization problem. The function you are minimizing is the sum of squared errors. The variables are the slope and the intercept. The derivative of the error with respect to each variable tells you how to nudge that variable to reduce the error.

You do not need to solve this by hand. Libraries do it for you. But the underlying mechanism is gradient descent: compute the gradient, take a step opposite the gradient, repeat.

## See it run

Here is a tiny gradient descent implementation that finds the minimum of `f(x) = (x - 3)^2`. The minimum is at `x = 3`.

```python runnable
def f(x):
    return (x - 3) ** 2

def derivative_f(x):
    return 2 * (x - 3)

x = 0.0          # starting point
learning_rate = 0.1

for i in range(20):
    grad = derivative_f(x)
    x = x - learning_rate * grad
    print(f"Step {i+1}: x = {x:.4f}, f(x) = {f(x):.4f}")

print("Final x:", x)
```

*What just happened:* The function `f(x) = (x - 3)^2` has its minimum at `x = 3`. Starting from `x = 0`, the gradient descent algorithm took 20 steps, each time moving opposite the gradient (which points toward the minimum). The learning rate of 0.1 controlled the step size. By step 20, `x` had converged to approximately 3.0000, and `f(x)` was essentially zero.

## For builders

Optimization is not only for math class. It is the engine behind much of the software you write and use.

- **Machine learning** - Training any model is optimization. The model has parameters. The loss function measures error. Gradient descent (or a variant) finds the parameters that minimize the loss.
- **Hyperparameter tuning** - Choosing the learning rate, the batch size, the number of layers: all of these are optimization problems. You want the settings that produce the best validation score.
- **Performance tuning** - Finding the fastest configuration for a database query, the optimal cache size, or the best thread pool size: all optimization.
- **Operations research** - Scheduling, routing, inventory management: all about finding the best allocation of limited resources.
- **Game AI** - Pathfinding in a game is often optimization: find the path with the lowest cost, where cost might be distance, time, or risk.

> The key insight: optimization is the search for the best outcome. Calculus makes it systematic by showing that the best outcome happens where the derivative is zero. Follow the gradient, and you will find a peak or a valley. The challenge is knowing whether it is the best one.

## What we have built

- A **derivative** is the rate of change at a single point.
- A **critical point** is where the derivative is zero - a candidate for a maximum or minimum.
- A **local maximum** is the best point in its neighborhood. A **global maximum** is the best point overall.
- **Gradient descent** finds a minimum by stepping opposite the gradient, repeated until convergence.
- The **learning rate** controls step size. Too small is slow. Too large is unstable.
- In code, gradient descent trains neural networks, fits models, and tunes hyperparameters.

A quick check before you move on:

```quiz
[
  {
    "q": "What is a critical point of a function?",
    "choices": ["A point where the function is undefined", "A point where the derivative is zero", "A point where the function reaches its global maximum", "A point where the function is increasing fastest"],
    "answer": 1,
    "explain": "A critical point is a point where the derivative is zero (or undefined). At a smooth maximum or minimum, the slope is flat, so the derivative is zero. Not every critical point is a maximum or minimum, but every smooth maximum or minimum is a critical point."
  },
  {
    "q": "In gradient descent, what does the learning rate control?",
    "choices": ["The number of iterations", "The size of each step taken in the direction opposite the gradient", "The initial value of the parameters", "The dimensionality of the problem"],
    "answer": 1,
    "explain": "The learning rate controls how large a step you take at each iteration. A small learning rate means slow but stable convergence. A large learning rate means faster progress but risk of overshooting the minimum or diverging."
  },
  {
    "q": "What is the difference between a local maximum and a global maximum?",
    "choices": ["A local maximum is higher than a global maximum", "A local maximum is the best point in its neighborhood; a global maximum is the best point overall", "A local maximum is for functions of one variable; a global maximum is for functions of many variables", "There is no difference; they are the same thing"],
    "answer": 1,
    "explain": "A local maximum is higher than all nearby points, but a distant peak may be taller. A global maximum is the highest point in the entire domain. Gradient descent is guaranteed to find a local maximum, but not necessarily the global one."
  }
]
```


---

# Integrals as the Total So Far

## The odometer that taught me everything

You are driving down a highway. The speedometer tells you how fast you are going right now. That is the derivative. The odometer tells you how far you have traveled in total. That is the integral.

The odometer does not show speed. It shows the accumulated total of all the speeds you have been driving, added up over time. If you drive 60 mph for one hour, you go 60 miles. If you drive 30 mph for the next hour, you go another 30 miles. The odometer reads 90.

But what if your speed is changing every second? What if you accelerate, brake, and accelerate again? The odometer still works. It adds up all the tiny distances you traveled in each tiny slice of time. The integral is the mathematical name for that adding-up process.

## From slices to total

Suppose you want to know the area under a curve. The curve might represent speed over time, profit over quantity, or probability over outcomes. The area is the total.

You can approximate the area by slicing the region into thin rectangles and adding up their areas. Each rectangle has a width (a small slice of the x-axis) and a height (the value of the function at that slice). The area of one rectangle is width times height. The total area is the sum of all the rectangles.

```
Area ≈ sum of (width * height) for each slice
```

As the slices get thinner, the approximation gets better. In the limit, as the width of each slice approaches zero, the sum becomes an **integral**.

```
Integral = limit of the sum as slice width -> 0
```

That limit is what the integral symbol means. The long S shape is a sum. The little numbers at the bottom and top say where to start and stop adding.

## The area under a curve

Consider the function `f(x) = x`. This is a straight line from the origin at a 45 degree angle. The area under this line from `x = 0` to `x = 2` is a triangle with base 2 and height 2.

```
Area = (1/2) * base * height = (1/2) * 2 * 2 = 2
```

The integral gives the same answer:

```
integral from 0 to 2 of x dx = [x^2 / 2] from 0 to 2 = (4 / 2) - (0 / 2) = 2
```

The antiderivative of `x` is `x^2 / 2`. Evaluate it at the upper limit and subtract the value at the lower limit. That is the **fundamental theorem of calculus**: integration and differentiation are inverse operations.

## Expected value: probability as an integral

In [Probability & Statistics](/guides/probability-and-statistics) you learned that expected value is the sum of each value times its probability. For a continuous random variable, the sum becomes an integral.

```
Expected value = integral of (x * probability_density(x)) dx
```

The integral adds up the value `x` weighted by how likely it is, across all possible values. The result is the long-run average you would expect if you repeated the experiment infinitely many times.

This is the same idea as the odometer. Instead of adding up tiny distances, you are adding up tiny probabilities weighted by their outcomes. The integral is the accumulation machine. It works for distance, for area, for probability, and for anything else that adds up.

## Riemann sums: the integral made concrete

A **Riemann sum** is a way to approximate an integral by slicing the region into rectangles and adding their areas. It is the bridge between the intuitive "add up thin slices" idea and the formal integral.

```
Riemann sum = sum of f(x_i) * delta_x for i = 1 to n
```

`x_i` is the position of the i-th slice. `delta_x` is the width of the slice. `f(x_i)` is the height. As `n` gets larger and `delta_x` gets smaller, the Riemann sum approaches the true integral.

This is exactly what a computer does when it computes an integral numerically. It cannot evaluate the limit symbolically, so it chops the region into many thin slices and adds them up. For most practical purposes, a few thousand slices are enough.

## See it run

Here is a Riemann sum approximation of the integral of `x` from 0 to 2, compared to the exact answer.

```python runnable
def f(x):
    return x

def riemann_sum(a, b, n):
    width = (b - a) / n
    total = 0
    for i in range(n):
        x = a + i * width
        total = total + f(x) * width
    return total

exact = 2.0  # integral of x from 0 to 2 is 2
for n in [10, 100, 1000, 10000]:
    approx = riemann_sum(0, 2, n)
    print(f"n = {n:5d}: approx = {approx:.6f}, error = {abs(approx - exact):.6f}")
```

*What just happened:* The `riemann_sum` function sliced the interval from 0 to 2 into `n` equal pieces, evaluated `f(x) = x` at the left edge of each piece, and added up the areas of the resulting rectangles. With `n = 10`, the approximation was 1.800000, with an error of 0.200000. With `n = 10000`, the approximation was 1.999900, with an error of 0.000100. As the number of slices increased, the approximation converged to the exact answer of 2.

## For builders

Integrals are not only for math class. They are the way you compute totals from rates.

- **Profiling and monitoring** - If you have a graph of requests per second over time, the integral of that graph is the total number of requests. Integrating a latency distribution gives you the total time spent waiting.
- **Physics simulations** - If you know the acceleration of an object, integrating once gives velocity, and integrating again gives position. Game engines and physics simulators do this every frame.
- **Probability and statistics** - The expected value of a continuous random variable is an integral. The cumulative distribution function is the integral of the probability density function.
- **Machine learning** - The loss during training is often integrated over time to produce a learning curve. The area under the precision-recall curve is the average precision.
- **Signal processing** - Filters that smooth or integrate a signal are computing running integrals. A moving average is a discrete approximation of an integral.

> The key insight: an integral is "the total so far." It adds up a million tiny pieces to compute a whole. The derivative breaks a whole into rates. The integral builds a whole from rates. They are two sides of the same coin.

## What we have built

- An **integral** adds up a function over an interval to compute a total.
- A **Riemann sum** approximates an integral by slicing the region into thin rectangles and adding their areas.
- The **fundamental theorem of calculus** says integration and differentiation are inverse operations.
- The **expected value** of a continuous random variable is an integral of `x` times the probability density.
- In code, numerical integration uses Riemann sums with many thin slices.
- Integrals compute total distance from speed, total requests from rate, and total probability from density.

A quick check before you go:

```quiz
[
  {
    "q": "What does an integral compute?",
    "choices": ["The rate of change at a single point", "The total accumulation of a function over an interval", "The maximum value of a function", "The slope of a curve"],
    "answer": 1,
    "explain": "An integral adds up a function over an interval to compute a total. For speed over time, it gives total distance. For probability density, it gives total probability. For requests per second, it gives total requests."
  },
  {
    "q": "What is a Riemann sum?",
    "choices": ["A method for solving differential equations", "An approximation of an integral by slicing into thin rectangles and adding their areas", "A type of derivative", "A way to find the maximum of a function"],
    "answer": 1,
    "explain": "A Riemann sum approximates an integral by dividing the interval into slices, computing the area of a rectangle for each slice, and summing them. As the slices get thinner, the approximation approaches the true integral."
  },
  {
    "q": "How are integration and differentiation related?",
    "choices": ["They are unrelated operations", "They are inverse operations: integrating a derivative gives back the original function (up to a constant)", "Integration is always harder than differentiation", "Differentiation is only for polynomials"],
    "answer": 1,
    "explain": "The fundamental theorem of calculus states that integration and differentiation are inverse operations. If you differentiate a function and then integrate the result, you get back the original function (plus a constant). This is why antiderivatives are useful for computing definite integrals."
  }
]
```
