The Rate of Change and Tangent Slopes hero
Lesson 1Derivatives and Sensitivity

The Rate of Change and Tangent Slopes

Master single-variable calculus foundations, transition from intuitive sensitivity to formal derivatives, and apply power rules to tangent slopes.

In Topic 1, we established how to measure a neural network's mistakes. When our 2-layer network made a confident prediction on the screenplay for 'Die Hard in Space' (y^=0.95\hat{y} = 0.95) and the movie flopped at the box office (y=0y = 0), we converted that mistake into a precise scalar penalty using Mean Squared Error:

L=(y−y^)2=(0−0.95)2=0.9025L = (y - \hat{y})^2 = (0 - 0.95)^2 = \mathbf{0.9025}

We also computed the derivative of the loss with respect to the output prediction:

∂L∂y^=2(y^−y)=2(0.95−0)=+1.90\frac{\partial L}{\partial \hat{y}} = 2(\hat{y} - y) = 2(0.95 - 0) = \mathbf{+1.90}

This positive value (+1.90+1.90) provided an unambiguous signal: because our prediction was too high (y^>y\hat{y} > y), the loss sits on an uphill slope. To reduce the loss penalty, the output prediction must decrease.

[ Forward Pass: Prediction ] ──► [ Loss Evaluation ] ──► [ Output Loss Sensitivity ]
        y_hat = 0.95                 L = 0.9025                 dL/dy_hat = +1.90

However, in a neural network, we cannot directly dial the output prediction y^\hat{y} by hand. The prediction is computed by combining internal weights (ww), biases (bb), and non-linear activation functions (σ(z)\sigma(z)).

If we want to adjust an internal parameter—such as the Action genre weight (w1w_1) in our greenlight predictor—we must answer a fundamental question:

"If we nudge an internal parameter by a microscopic amount, how much will the output change?"\text{\textit{"If we nudge an internal parameter by a microscopic amount, how much will the output change?"}}

To answer this question, we must understand the mathematical engine of learning: single-variable differential calculus.

[ Topic 1: Loss Functions ] ──► [ Topic 2: Derivatives ] ──► [ Topic 3: Gradients ] ──► [ Topic 4: Backpropagation ]
      L = (y - y_hat)^2               dy/dx & f'(x)                 grad_w L = dL/dw             Chain Rule Across Layers

In this topic, we build the calculus toolkit from first principles: measuring rates of change, collapsing secant lines into tangent slopes via limits, transitioning from intuitive sensitivity to formal derivative notation, and mastering foundational algebraic power rules.



Measure rates of change using secant lines and tangent slopes to quantify how responsive dependent outputs are to variations in independent inputs.

Before analyzing complex curves or neural networks, we must formalize how one quantity responds to changes in another.


Defining Rate of Change: Linear Relationships

In the simplest mathematical systems, the relationship between two variables is a straight line.

Consider a simple hourly wage calculator:

y=15xy = 15x

where xx represents the number of hours worked, and yy represents total earnings in dollars.

If you work x1=2x_1 = 2 hours, you earn y1=$30y_1 = \char36 30. If you increase your hours to x2=5x_2 = 5 hours, you earn y2=$75y_2 = \char36 75.

The rate of change measures the ratio of the change in output (Δy\Delta y) to the change in input (Δx\Delta x):

ΔyΔx=y2−y1x2−x1=75−305−2=453=15 dollars per hour\frac{\Delta y}{\Delta x} = \frac{y_2 - y_1}{x_2 - x_1} = \frac{75 - 30}{5 - 2} = \frac{45}{3} = \mathbf{15 \text{ dollars per hour}}

The symbol Δ\Delta (the Greek letter delta) denotes a macroscopic, measurable difference:

Δx=x2−x1,Δy=y2−y1\Delta x = x_2 - x_1, \qquad \Delta y = y_2 - y_1

In any linear equation y=mx+cy = mx + c, the rate of change is the slope (mm).

Whether you compare 22 hours to 55 hours, or 100100 hours to 100.5100.5 hours:

ΔyΔx=15(100.5)−15(100)100.5−100=1507.5−15000.5=7.50.5=15\frac{\Delta y}{\Delta x} = \frac{15(100.5) - 15(100)}{100.5 - 100} = \frac{1507.5 - 1500}{0.5} = \frac{7.5}{0.5} = \mathbf{15}

The slope is constant everywhere. Every additional hour worked always produces exactly $15\char36 15.

Dual-Track Bridge: From Wages to Neural Sensitivity

  • Track 1 (Underlying Mechanism): In our hourly wage calculator (y=15xy = 15x), the sensitivity of earnings to hours worked is a constant ΔyΔx=15 dollars/hour\frac{\Delta y}{\Delta x} = 15\text{ dollars/hour}. Every unit change in xx produces an identical $15\char36 15 change in yy.
  • Track 2 (Applied Concept): In a neural network pre-activation term z=w1x1+bz = w_1 x_1 + b, if the Action genre feature is present (x1=1.0x_1 = 1.0) for 'Die Hard in Space', the pre-activation becomes z(w1)=1.0⋅w1+bz(w_1) = 1.0 \cdot w_1 + b. The rate of change with respect to the Action weight is a constant ΔzΔw1=x1=1.0\frac{\Delta z}{\Delta w_1} = x_1 = 1.0. Nudging w1w_1 by +0.2+0.2 nudges the linear sum zz by +0.2+0.2.

Non-Linear Curves: The Vanishing Single Slope

In machine learning, systems are rarely straight lines. Neural networks rely on non-linear activation curves (such as Sigmoid σ(z)\sigma(z)) and quadratic error surfaces (L=(y^−y)2L = (\hat{y} - y)^2).

Consider the simple quadratic curve:

f(x)=x2f(x) = x^2

Let's examine the rate of change across different intervals along this curve:

  • From x=0x = 0 to x=1x = 1: ΔyΔx=f(1)−f(0)1−0=1−01=1.0\frac{\Delta y}{\Delta x} = \frac{f(1) - f(0)}{1 - 0} = \frac{1 - 0}{1} = \mathbf{1.0}

  • From x=1x = 1 to x=2x = 2: ΔyΔx=f(2)−f(1)2−1=4−11=3.0\frac{\Delta y}{\Delta x} = \frac{f(2) - f(1)}{2 - 1} = \frac{4 - 1}{1} = \mathbf{3.0}

  • From x=4x = 4 to x=5x = 5: ΔyΔx=f(5)−f(4)5−4=25−161=9.0\frac{\Delta y}{\Delta x} = \frac{f(5) - f(4)}{5 - 4} = \frac{25 - 16}{1} = \mathbf{9.0}

Interval [0, 1]:  Average Slope = 1.0  (Gentle rise)
Interval [1, 2]:  Average Slope = 3.0  (Moderate rise)
Interval [4, 5]:  Average Slope = 9.0  (Steep wall)

There is no single number that describes the slope of f(x)=x2f(x) = x^2. The rate of change changes continuously depending on where you stand on the curve.


Secant Lines: Average Rate of Change over an Interval

A line drawn through two distinct points on a curve is called a Secant Line.

For any function f(x)f(x), if we start at an operating point xx and move forward by an interval Δx\Delta x, the secant line connects the points (x,f(x))(x, f(x)) and (x+Δx,f(x+Δx))(x + \Delta x, f(x + \Delta x)).

The slope of the secant line represents the average rate of change over that interval:

msecant=ΔyΔx=f(x+Δx)−f(x)Δxm_{\text{secant}} = \frac{\Delta y}{\Delta x} = \frac{f(x + \Delta x) - f(x)}{\Delta x}

Let's fix our starting point at x=2x = 2 (where f(2)=4f(2) = 4) and calculate the secant slope as we shrink the interval Δx\Delta x:

Starting Point (xx)Interval Step (Δx\Delta x)Second Point (x+Δxx + \Delta x)Function Value f(x+Δx)f(x + \Delta x)Output Change (Δy\Delta y)Secant Slope (ΔyΔx\frac{\Delta y}{\Delta x})
2.02.01.01.03.03.0(3.0)2=9.0(3.0)^2 = 9.09.0−4.0=5.09.0 - 4.0 = 5.05.01.0=5.0\frac{5.0}{1.0} = \mathbf{5.0}
2.02.00.50.52.52.5(2.5)2=6.25(2.5)^2 = 6.256.25−4.0=2.256.25 - 4.0 = 2.252.250.5=4.5\frac{2.25}{0.5} = \mathbf{4.5}
2.02.00.10.12.12.1(2.1)2=4.41(2.1)^2 = 4.414.41−4.0=0.414.41 - 4.0 = 0.410.410.1=4.1\frac{0.41}{0.1} = \mathbf{4.1}
2.02.00.010.012.012.01(2.01)2=4.0401(2.01)^2 = 4.04014.0401−4.0=0.04014.0401 - 4.0 = 0.04010.04010.01=4.01\frac{0.0401}{0.01} = \mathbf{4.01}
2.02.00.0010.0012.0012.001(2.001)2=4.004001(2.001)^2 = 4.0040014.004001−4.0=0.0040014.004001 - 4.0 = 0.0040010.0040010.001=4.001\frac{0.004001}{0.001} = \mathbf{4.001}

Observe the sequence of secant slopes as Δx\Delta x shrinks toward zero:

5.0⟶4.5⟶4.1⟶4.01⟶4.001⟶4.05.0 \longrightarrow 4.5 \longrightarrow 4.1 \longrightarrow 4.01 \longrightarrow 4.001 \longrightarrow \mathbf{4.0}

As the step size Δx\Delta x becomes microscopic, the average slope converges cleanly toward a single exact number: 4.04.0.


Tangent Lines: Instantaneous Slope at an Operating Point

When the step size Δx\Delta x shrinks toward zero, the second point slides along the curve until it merges with the starting point.

The secant line pivots until it grazes the curve at that single point without cutting through it. This grazing line is the Tangent Line.

    Secant Line (Two points)                Tangent Line (One point)
         y                                       y
         |         * (x+Δx, f(x+Δx))             |          .
         |       /                               |        /
         |     * (x, f(x))                       |     * (x, f(x))  [Slope = 4.0]
         |    /                                  |    /
         +─────────────── x                      +─────────────── x
  • The Secant Line: Measures the average rate of change between two separated points across a finite distance Δx\Delta x.
  • The Tangent Line: Measures the instantaneous rate of change at a single operating point (x,f(x))(x, f(x)).

The slope of this tangent line tells us the exact sensitivity of the curve at that precise coordinate. In Lesson 2, we formalize this geometric convergence into the foundational equation of differential calculus.


Previous
Prediction Error & Loss Functions In Practice