
The Composite Function Chain Rule
Apply composite chain rules to attribute error through output and hidden layers, computing backward gradients and toy parameter updates on paper.
In Topic 1, we measured prediction error on the screenplay for 'Die Hard in Space' using Mean Squared Error:
In Topic 2, we established single-variable calculus to compute the instantaneous sensitivity of an isolated operation. In Topic 3, we used the freezing principle to calculate partial derivatives for a single linear neuron, assembling them into the gradient vector to determine the direction of steepest error descent.
[ Module 1: The Forward Pass ]
Topic 1: Feature Vectors & Dot Products (x, w · x)
Topic 2: Affine Sums & Activations (z = w · x + b, a = σ(z))
Topic 3: Parallel Layers & Width (z = W x + b, a = σ(z))
Topic 4: Multi-Layer Perceptrons & Depth (a^(1) = ReLU(W^(1)x + b^(1)), a^(2) = σ(W^(2)a^(1) + b^(2)))
│
▼
[ Module 2: The Backward Pass ]
Topic 1: Prediction Error & Squared Loss (L = (y - y_hat)^2, ∂L/∂y_hat = 2(y_hat - y))
Topic 2: Single-Variable Derivatives (dy/dx, tangent slopes, power/sum rules)
Topic 3: Multivariable Gradients (∂L/∂w_i, ∇_w L, steepest descent -∇L)
Topic 4: The Chain Rule & Multilayer Backpropagation ◄─── [THIS TOPIC]
However, real deep neural networks never connect raw inputs directly to output losses through a single linear sum.
In a multi-layer network, parameters are buried deep inside composite functional hierarchies. A first-layer weight does not touch the loss directly. Instead, it alters a hidden pre-activation , which passes through a non-linear activation function , which feeds into an output pre-activation , which passes through an output activation , which finally produces the loss penalty :
To calculate the gradient of the loss with respect to buried internal weights, we cannot use simple single-layer partial derivatives. We need a systematic mathematical engine capable of propagating sensitivity backward through chained functional transformations.
That engine is The Chain Rule, and its algorithmic execution across neural networks is Backpropagation.
Derive the single-variable chain rule to multiply rates of change across composite functions and propagate sensitivity through chained operations.
In ordinary single-variable calculus, a function maps an input directly to an output .
In a chained system, functions are composed into pipelines: an input passes into an intermediate function , and the resulting intermediate value passes into an outer function :
When the input is nudged by a tiny amount , how much does the final output change?
The Rate Multiplication Principle
To understand how sensitivity flows through a chain of operations, consider how local rates of change compound.
Suppose:
- When input changes, intermediate variable changes 3 times as fast (rate of change ).
- When intermediate variable changes, output variable changes 2 times as fast (rate of change ).
If you nudge upward by unit:
- Intermediate increases by units ().
- Because increased by units, output increases by units ().
The overall rate of change of with respect to is the product of the individual rates:
Rates of change across composite functional dependencies do not add—they multiply.
Dual-Track Bridge: Mechanical Gear Trains
To anchor this rate multiplication principle in a physical system, examine a three-gear mechanical transmission.
[ Gear A (Input x) ] ──(3:1 Ratio)──► [ Gear B (Intermediate u) ] ──(2:1 Ratio)──► [ Gear C (Output y) ]
30 Teeth 10 Teeth 5 Teeth
-
Track 1 (Underlying Mechanism):
- Gear A (input ) has 30 teeth and meshes with Gear B (intermediate ), which has 10 teeth. For every 1 full turn of Gear A, Gear B rotates 3 full turns:
- Gear B is mounted on a shaft that meshes with Gear C (output ), which has 5 teeth. For every 1 full turn of Gear B, Gear C rotates 2 full turns:
- When you rotate Gear A by 1 turn, Gear C rotates by turns:
-
Track 2 (Applied Neural Architecture):
- In a neural network, Gear A represents a feature measurement , Gear B represents the hidden linear sum , and Gear C represents the final output prediction .
- If nudging changes at a rate of , and nudging changes at a rate of , the compound sensitivity of the prediction to the input is .
Formal Definition and Limit Derivation
Mathematically, the single-variable Chain Rule states that for composite function where and , if is differentiable at and is differentiable at , then:
Step-by-Step Limit Proof
We define the derivative of with respect to using the formal limit definition:
Let be the resulting change in the intermediate variable. For non-zero , we multiply and divide the difference quotient by :
Because is continuous, as , the change . Applying the limit product law:
Evaluating both limits yields the Chain Rule formula:
Step-by-Step Algebraic Verification:
To verify that the chain rule produces exact analytical derivatives, let us differentiate the composite expression using two independent mathematical methods.
Method A: Direct Polynomial Expansion
First, expand the polynomial algebraically before differentiating:
Apply the standard power and sum rules term-by-term:
Method B: Chain Rule Decomposition
Decompose the expression into outer and inner functions:
- Inner function:
- Outer function:
Differentiate each component separately:
- Derivative of outer function with respect to :
- Derivative of inner function with respect to :
Multiply the derivatives together according to the chain rule:
Substitute the original inner expression back into the formula:
Numerical Evaluation at
- Under Method A:
- Under Method B:
Both methods yield identical derivatives (). The chain rule allows us to differentiate deeply nested functions without expanding algebraic polynomials.