
The Rate of Change and Tangent Slopes
Master single-variable calculus foundations, transition from intuitive sensitivity to formal derivatives, and apply power rules to tangent slopes.
In Topic 1, we established how to measure a neural network's mistakes. When our 2-layer network made a confident prediction on the screenplay for 'Die Hard in Space' () and the movie flopped at the box office (), we converted that mistake into a precise scalar penalty using Mean Squared Error:
We also computed the derivative of the loss with respect to the output prediction:
This positive value () provided an unambiguous signal: because our prediction was too high (), the loss sits on an uphill slope. To reduce the loss penalty, the output prediction must decrease.
[ Forward Pass: Prediction ] ──► [ Loss Evaluation ] ──► [ Output Loss Sensitivity ]
y_hat = 0.95 L = 0.9025 dL/dy_hat = +1.90
However, in a neural network, we cannot directly dial the output prediction by hand. The prediction is computed by combining internal weights (), biases (), and non-linear activation functions ().
If we want to adjust an internal parameter—such as the Action genre weight () in our greenlight predictor—we must answer a fundamental question:
To answer this question, we must understand the mathematical engine of learning: single-variable differential calculus.
[ Topic 1: Loss Functions ] ──► [ Topic 2: Derivatives ] ──► [ Topic 3: Gradients ] ──► [ Topic 4: Backpropagation ]
L = (y - y_hat)^2 dy/dx & f'(x) grad_w L = dL/dw Chain Rule Across Layers
In this topic, we build the calculus toolkit from first principles: measuring rates of change, collapsing secant lines into tangent slopes via limits, transitioning from intuitive sensitivity to formal derivative notation, and mastering foundational algebraic power rules.
Measure rates of change using secant lines and tangent slopes to quantify how responsive dependent outputs are to variations in independent inputs.
Before analyzing complex curves or neural networks, we must formalize how one quantity responds to changes in another.
Defining Rate of Change: Linear Relationships
In the simplest mathematical systems, the relationship between two variables is a straight line.
Consider a simple hourly wage calculator:
where represents the number of hours worked, and represents total earnings in dollars.
If you work hours, you earn . If you increase your hours to hours, you earn .
The rate of change measures the ratio of the change in output () to the change in input ():
The symbol (the Greek letter delta) denotes a macroscopic, measurable difference:
In any linear equation , the rate of change is the slope ().
Whether you compare hours to hours, or hours to hours:
The slope is constant everywhere. Every additional hour worked always produces exactly .
Dual-Track Bridge: From Wages to Neural Sensitivity
- Track 1 (Underlying Mechanism): In our hourly wage calculator (), the sensitivity of earnings to hours worked is a constant . Every unit change in produces an identical change in .
- Track 2 (Applied Concept): In a neural network pre-activation term , if the Action genre feature is present () for 'Die Hard in Space', the pre-activation becomes . The rate of change with respect to the Action weight is a constant . Nudging by nudges the linear sum by .
Non-Linear Curves: The Vanishing Single Slope
In machine learning, systems are rarely straight lines. Neural networks rely on non-linear activation curves (such as Sigmoid ) and quadratic error surfaces ().
Consider the simple quadratic curve:
Let's examine the rate of change across different intervals along this curve:
-
From to :
-
From to :
-
From to :
Interval [0, 1]: Average Slope = 1.0 (Gentle rise)
Interval [1, 2]: Average Slope = 3.0 (Moderate rise)
Interval [4, 5]: Average Slope = 9.0 (Steep wall)
There is no single number that describes the slope of . The rate of change changes continuously depending on where you stand on the curve.
Secant Lines: Average Rate of Change over an Interval
A line drawn through two distinct points on a curve is called a Secant Line.
For any function , if we start at an operating point and move forward by an interval , the secant line connects the points and .
The slope of the secant line represents the average rate of change over that interval:
Let's fix our starting point at (where ) and calculate the secant slope as we shrink the interval :
| Starting Point () | Interval Step () | Second Point () | Function Value | Output Change () | Secant Slope () |
|---|---|---|---|---|---|
Observe the sequence of secant slopes as shrinks toward zero:
As the step size becomes microscopic, the average slope converges cleanly toward a single exact number: .
Tangent Lines: Instantaneous Slope at an Operating Point
When the step size shrinks toward zero, the second point slides along the curve until it merges with the starting point.
The secant line pivots until it grazes the curve at that single point without cutting through it. This grazing line is the Tangent Line.
Secant Line (Two points) Tangent Line (One point)
y y
| * (x+Δx, f(x+Δx)) | .
| / | /
| * (x, f(x)) | * (x, f(x)) [Slope = 4.0]
| / | /
+─────────────── x +─────────────── x
- The Secant Line: Measures the average rate of change between two separated points across a finite distance .
- The Tangent Line: Measures the instantaneous rate of change at a single operating point .
The slope of this tangent line tells us the exact sensitivity of the curve at that precise coordinate. In Lesson 2, we formalize this geometric convergence into the foundational equation of differential calculus.