The Composite Function Chain Rule hero
Lesson 1The Chain Rule and Backpropagation

The Composite Function Chain Rule

Apply composite chain rules to attribute error through output and hidden layers, computing backward gradients and toy parameter updates on paper.

In Topic 1, we measured prediction error on the screenplay for 'Die Hard in Space' using Mean Squared Error:

L=(y−y^)2=(0−0.9820)2≈0.9644L = (y - \hat{y})^2 = (0 - 0.9820)^2 \approx \mathbf{0.9644}

In Topic 2, we established single-variable calculus to compute the instantaneous sensitivity dydx\frac{dy}{dx} of an isolated operation. In Topic 3, we used the freezing principle to calculate partial derivatives ∂L∂wi\frac{\partial L}{\partial w_i} for a single linear neuron, assembling them into the gradient vector ∇L\nabla L to determine the direction of steepest error descent.

[ Module 1: The Forward Pass ]
  Topic 1: Feature Vectors & Dot Products (x, w · x)
  Topic 2: Affine Sums & Activations (z = w · x + b, a = σ(z))
  Topic 3: Parallel Layers & Width (z = W x + b, a = σ(z))
  Topic 4: Multi-Layer Perceptrons & Depth (a^(1) = ReLU(W^(1)x + b^(1)), a^(2) = σ(W^(2)a^(1) + b^(2)))
                                │
                                ▼
[ Module 2: The Backward Pass ]
  Topic 1: Prediction Error & Squared Loss (L = (y - y_hat)^2, ∂L/∂y_hat = 2(y_hat - y))
  Topic 2: Single-Variable Derivatives (dy/dx, tangent slopes, power/sum rules)
  Topic 3: Multivariable Gradients (∂L/∂w_i, ∇_w L, steepest descent -∇L)
  Topic 4: The Chain Rule & Multilayer Backpropagation ◄─── [THIS TOPIC]

However, real deep neural networks never connect raw inputs directly to output losses through a single linear sum.

In a multi-layer network, parameters are buried deep inside composite functional hierarchies. A first-layer weight Wij(1)W^{(1)}_{ij} does not touch the loss directly. Instead, it alters a hidden pre-activation zi(1)z^{(1)}_i, which passes through a non-linear activation function ai(1)a^{(1)}_i, which feeds into an output pre-activation z(2)z^{(2)}, which passes through an output activation y^=a(2)\hat{y} = a^{(2)}, which finally produces the loss penalty LL:

x→W(1),b(1)z(1)→fa(1)→W(2),b(2)z(2)→σa(2)→LossLx \xrightarrow{W^{(1)}, b^{(1)}} z^{(1)} \xrightarrow{f} a^{(1)} \xrightarrow{W^{(2)}, b^{(2)}} z^{(2)} \xrightarrow{\sigma} a^{(2)} \xrightarrow{\text{Loss}} L

To calculate the gradient of the loss with respect to buried internal weights, we cannot use simple single-layer partial derivatives. We need a systematic mathematical engine capable of propagating sensitivity backward through chained functional transformations.

That engine is The Chain Rule, and its algorithmic execution across neural networks is Backpropagation.



Derive the single-variable chain rule to multiply rates of change across composite functions and propagate sensitivity through chained operations.

In ordinary single-variable calculus, a function maps an input xx directly to an output y=f(x)y = f(x).

In a chained system, functions are composed into pipelines: an input xx passes into an intermediate function u=g(x)u = g(x), and the resulting intermediate value uu passes into an outer function y=f(u)y = f(u):

y=(f∘g)(x)=f(g(x))y = (f \circ g)(x) = f(g(x))

When the input xx is nudged by a tiny amount Δx\Delta x, how much does the final output yy change?


The Rate Multiplication Principle

To understand how sensitivity flows through a chain of operations, consider how local rates of change compound.

Suppose:

  1. When input xx changes, intermediate variable uu changes 3 times as fast (rate of change dudx=3\frac{du}{dx} = 3).
  2. When intermediate variable uu changes, output variable yy changes 2 times as fast (rate of change dydu=2\frac{dy}{du} = 2).

If you nudge xx upward by +1+1 unit:

  • Intermediate uu increases by +3+3 units (1×3=31 \times 3 = 3).
  • Because uu increased by +3+3 units, output yy increases by +6+6 units (3×2=63 \times 2 = 6).

The overall rate of change of yy with respect to xx is the product of the individual rates:

dydx=dydu⋅dudx=2×3=6\frac{dy}{dx} = \frac{dy}{du} \cdot \frac{du}{dx} = 2 \times 3 = \mathbf{6}

Rates of change across composite functional dependencies do not add—they multiply.


Dual-Track Bridge: Mechanical Gear Trains

To anchor this rate multiplication principle in a physical system, examine a three-gear mechanical transmission.

[ Gear A (Input x) ] ──(3:1 Ratio)──► [ Gear B (Intermediate u) ] ──(2:1 Ratio)──► [ Gear C (Output y) ]
     30 Teeth                              10 Teeth                              5 Teeth
  • Track 1 (Underlying Mechanism):

    • Gear A (input xx) has 30 teeth and meshes with Gear B (intermediate uu), which has 10 teeth. For every 1 full turn of Gear A, Gear B rotates 3 full turns: dudx=3\frac{du}{dx} = 3
    • Gear B is mounted on a shaft that meshes with Gear C (output yy), which has 5 teeth. For every 1 full turn of Gear B, Gear C rotates 2 full turns: dydu=2\frac{dy}{du} = 2
    • When you rotate Gear A by 1 turn, Gear C rotates by 3×2=63 \times 2 = 6 turns: dydx=dydu⋅dudx=2×3=6\frac{dy}{dx} = \frac{dy}{du} \cdot \frac{du}{dx} = 2 \times 3 = \mathbf{6}
  • Track 2 (Applied Neural Architecture):

    • In a neural network, Gear A represents a feature measurement x1x_1, Gear B represents the hidden linear sum z1(1)z_1^{(1)}, and Gear C represents the final output prediction y^\hat{y}.
    • If nudging x1x_1 changes z1(1)z_1^{(1)} at a rate of +3.0+3.0, and nudging z1(1)z_1^{(1)} changes y^\hat{y} at a rate of +0.5+0.5, the compound sensitivity of the prediction to the input is (+3.0)×(+0.5)=+1.5(+3.0) \times (+0.5) = +1.5.

Formal Definition and Limit Derivation

Mathematically, the single-variable Chain Rule states that for composite function y=f(g(x))y = f(g(x)) where u=g(x)u = g(x) and y=f(u)y = f(u), if gg is differentiable at xx and ff is differentiable at g(x)g(x), then:

dydx=dydu⋅dudx\frac{dy}{dx} = \frac{dy}{du} \cdot \frac{du}{dx}

Step-by-Step Limit Proof

We define the derivative of yy with respect to xx using the formal limit definition:

dydx=lim⁡Δx→0ΔyΔx\frac{dy}{dx} = \lim_{\Delta x \to 0} \frac{\Delta y}{\Delta x}

Let Δu=g(x+Δx)−g(x)\Delta u = g(x + \Delta x) - g(x) be the resulting change in the intermediate variable. For non-zero Δu\Delta u, we multiply and divide the difference quotient by Δu\Delta u:

ΔyΔx=ΔyΔu⋅ΔuΔx\frac{\Delta y}{\Delta x} = \frac{\Delta y}{\Delta u} \cdot \frac{\Delta u}{\Delta x}

Because gg is continuous, as Δx→0\Delta x \to 0, the change Δu→0\Delta u \to 0. Applying the limit product law:

lim⁡Δx→0ΔyΔx=(lim⁡Δu→0ΔyΔu)⋅(lim⁡Δx→0ΔuΔx)\lim_{\Delta x \to 0} \frac{\Delta y}{\Delta x} = \left( \lim_{\Delta u \to 0} \frac{\Delta y}{\Delta u} \right) \cdot \left( \lim_{\Delta x \to 0} \frac{\Delta u}{\Delta x} \right)

Evaluating both limits yields the Chain Rule formula:

dydx=dydu⋅dudx\frac{dy}{dx} = \frac{dy}{du} \cdot \frac{du}{dx}

Step-by-Step Algebraic Verification: y=(3x+1)2y = (3x + 1)^2

To verify that the chain rule produces exact analytical derivatives, let us differentiate the composite expression y=(3x+1)2y = (3x + 1)^2 using two independent mathematical methods.

Method A: Direct Polynomial Expansion

First, expand the polynomial algebraically before differentiating:

y=(3x+1)2=(3x+1)(3x+1)=9x2+6x+1y = (3x + 1)^2 = (3x + 1)(3x + 1) = 9x^2 + 6x + 1

Apply the standard power and sum rules term-by-term:

dydx=ddx[9x2]+ddx[6x]+ddx[1]=18x+6\frac{dy}{dx} = \frac{d}{dx}[9x^2] + \frac{d}{dx}[6x] + \frac{d}{dx}[1] = 18x + 6

Method B: Chain Rule Decomposition

Decompose the expression into outer and inner functions:

  • Inner function: u=g(x)=3x+1u = g(x) = 3x + 1
  • Outer function: y=f(u)=u2y = f(u) = u^2

Differentiate each component separately:

  1. Derivative of outer function with respect to uu: dydu=ddu[u2]=2u\frac{dy}{du} = \frac{d}{du}[u^2] = 2u
  2. Derivative of inner function with respect to xx: dudx=ddx[3x+1]=3\frac{du}{dx} = \frac{d}{dx}[3x + 1] = 3

Multiply the derivatives together according to the chain rule:

dydx=dydu⋅dudx=(2u)⋅(3)=6u\frac{dy}{dx} = \frac{dy}{du} \cdot \frac{du}{dx} = (2u) \cdot (3) = 6u

Substitute the original inner expression u=3x+1u = 3x + 1 back into the formula:

dydx=6(3x+1)=18x+6\frac{dy}{dx} = 6(3x + 1) = 18x + 6

Numerical Evaluation at x=2x = 2

  • Under Method A: dydx=18(2)+6=36+6=42\frac{dy}{dx} = 18(2) + 6 = 36 + 6 = \mathbf{42}
  • Under Method B: u=3(2)+1=7  ⟹  dydx=6(7)=42u = 3(2) + 1 = 7 \implies \frac{dy}{dx} = 6(7) = \mathbf{42}

Both methods yield identical derivatives (18x+618x + 6). The chain rule allows us to differentiate deeply nested functions without expanding algebraic polynomials.


Previous
Partial Derivatives & Gradient Vectors In Practice