The Chain Rule & Backpropagation In Practice hero
LaboratoryThe Chain Rule and Backpropagation

The Chain Rule & Backpropagation In Practice

Master composite chain rules, output layer error attribution, hidden backpropagation deltas, and single-step parameter updates through manual calculations.

Practice active mathematical verification on paper before revealing the step-by-step solutions.


Part 1: Abstract Mechanics

Work through these 5 rapid-fire drill problems covering the core calculus of the backward pass.

Problem 1: Composite Single-Variable Chain Rule

Given the composite function y=u3y = u^3 where u=2x−4u = 2x - 4:

  • Derive the formula for dydx\frac{dy}{dx} in terms of xx using the chain rule.
  • Calculate the exact numerical derivative at x=3x = 3.
Reveal Solution
  1. Differentiate outer function: dydu=ddu[u3]=3u2\frac{dy}{du} = \frac{d}{du}[u^3] = 3u^2
  2. Differentiate inner function: dudx=ddx[2x−4]=2\frac{du}{dx} = \frac{d}{dx}[2x - 4] = 2
  3. Apply chain rule: dydx=dydu⋅dudx=(3u2)(2)=6u2=6(2x−4)2\frac{dy}{dx} = \frac{dy}{du} \cdot \frac{du}{dx} = (3u^2)(2) = 6u^2 = 6(2x - 4)^2
  4. Expanded form: 6(4x2−16x+16)=24x2−96x+966(4x^2 - 16x + 16) = \mathbf{24x^2 - 96x + 96}
  5. Evaluate at x=3x = 3: u=2(3)−4=2  ⟹  dydx=6(2)2=6(4)=24.0u = 2(3) - 4 = 2 \implies \frac{dy}{dx} = 6(2)^2 = 6(4) = \mathbf{24.0}

(Verification via expansion: y=(2x−4)3=8x3−48x2+96x−64  ⟹  dydx=24x2−96x+96y = (2x - 4)^3 = 8x^3 - 48x^2 + 96x - 64 \implies \frac{dy}{dx} = 24x^2 - 96x + 96. At x=3x=3: 24(9)−96(3)+96=216−288+96=24.024(9) - 96(3) + 96 = 216 - 288 + 96 = 24.0.)


Problem 2: Sigmoid Activation Derivative Calculation

A neuron produces a post-activation output a=σ(z)=0.80a = \sigma(z) = 0.80. Compute the exact instantaneous activation slope σ′(z)=σ(z)(1−σ(z))\sigma'(z) = \sigma(z)(1 - \sigma(z)).

Reveal Solution
σ′(z)=σ(z)(1−σ(z))=0.80×(1−0.80)=0.80×0.20=0.1600\sigma'(z) = \sigma(z)(1 - \sigma(z)) = 0.80 \times (1 - 0.80) = 0.80 \times 0.20 = \mathbf{0.1600}

Problem 3: Output Layer Error Delta δ(2)\delta^{(2)} Computation

An output neuron has loss derivative ∂L∂a(2)=1.60\frac{\partial L}{\partial a^{(2)}} = 1.60 and activation slope σ′(z(2))=0.25\sigma'(z^{(2)}) = 0.25. Compute the output error delta δ(2)=∂L∂a(2)⋅σ′(z(2))\delta^{(2)} = \frac{\partial L}{\partial a^{(2)}} \cdot \sigma'(z^{(2)}).

Reveal Solution
δ(2)=∂L∂a(2)⋅σ′(z(2))=1.60×0.25=0.4000\delta^{(2)} = \frac{\partial L}{\partial a^{(2)}} \cdot \sigma'(z^{(2)}) = 1.60 \times 0.25 = \mathbf{0.4000}

Problem 4: Hidden Layer Error Delta via Transposed Weights

An output layer error delta is δ(2)=0.40\delta^{(2)} = 0.40, connected via output weight W(2)=0.50W^{(2)} = 0.50 to a hidden neuron with activation slope σ′(z(1))=0.20\sigma'(z^{(1)}) = 0.20. Compute the hidden error delta δ(1)=((W(2))Tδ(2))⋅σ′(z(1))\delta^{(1)} = ((W^{(2)})^T \delta^{(2)}) \cdot \sigma'(z^{(1)}).

Reveal Solution
  1. Project error through transposed weight: (W(2))Tδ(2)=0.50×0.40=0.2000(W^{(2)})^T \delta^{(2)} = 0.50 \times 0.40 = 0.2000
  2. Modulate by hidden activation slope: δ(1)=0.2000×0.20=0.0400\delta^{(1)} = 0.2000 \times 0.20 = \mathbf{0.0400}

Problem 5: Single Parameter Update Step Calculation

A weight parameter has current value wold=0.60w_{\text{old}} = 0.60, computed gradient ∂L∂w=0.40\frac{\partial L}{\partial w} = 0.40, and step size factor η=0.10\eta = 0.10. Calculate the updated weight wnew=wold−η∂L∂ww_{\text{new}} = w_{\text{old}} - \eta \frac{\partial L}{\partial w}.

Reveal Solution
wnew=wold−η∂L∂w=0.60−(0.10×0.40)=0.60−0.04=0.5600w_{\text{new}} = w_{\text{old}} - \eta \frac{\partial L}{\partial w} = 0.60 - (0.10 \times 0.40) = 0.60 - 0.04 = \mathbf{0.5600}

Part 2: Applied Scenario: The VC 2-Layer Backpropagation Trace

An automated venture capital risk model uses a 2→2→12 \to 2 \to 1 Multi-Layer Perceptron to predict whether an early-stage startup will succeed (y=1y = 1) or go bankrupt (y=0y = 0).

System Configuration:

  • Input Feature Vector: x=[1.02.0](Founding Team Experience)(Market Size)x = \begin{bmatrix} 1.0 \\ 2.0 \end{bmatrix} \begin{matrix} \text{(Founding Team Experience)} \\ \text{(Market Size)} \end{matrix}
  • Initial Layer 1 Parameters: W(1)=[0.50.50.50.5],b(1)=[0.00.0]W^{(1)} = \begin{bmatrix} 0.5 & 0.5 \\ 0.5 & 0.5 \end{bmatrix}, \qquad b^{(1)} = \begin{bmatrix} 0.0 \\ 0.0 \end{bmatrix}
  • Initial Layer 2 Parameters: W(2)=[0.50.5],b(2)=[0.0]W^{(2)} = \begin{bmatrix} 0.5 & 0.5 \end{bmatrix}, \qquad b^{(2)} = \begin{bmatrix} 0.0 \end{bmatrix}
  • Activation Functions: Standard Sigmoid activation σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}} across all hidden and output neurons.
  • Ground-Truth Target: y=0y = 0 (The startup went bankrupt).
  • Step Size (Learning Rate): η=0.50\eta = 0.50.

Problem 6: Complete 2-Layer VC Backward Pass Audit

  • Part A (Forward Pass & Initial Loss): Compute z(1),a(1),z(2),a(2)=y^z^{(1)}, a^{(1)}, z^{(2)}, a^{(2)} = \hat{y}, and initial MSE loss L1=(a(2)−y)2L_1 = (a^{(2)} - y)^2.
  • Part B (Output Layer Error Attribution): Compute ∂L∂a(2)\frac{\partial L}{\partial a^{(2)}}, σ′(z(2))\sigma'(z^{(2)}), output error delta δ(2)\delta^{(2)}, output weight gradient matrix ∂L∂W(2)\frac{\partial L}{\partial W^{(2)}}, and output bias gradient ∂L∂b(2)\frac{\partial L}{\partial b^{(2)}}.
  • Part C (Hidden Layer Backpropagation): Compute backpropagated error signal (W(2))Tδ(2)(W^{(2)})^T \delta^{(2)}, hidden activation slopes σ′(z(1))\sigma'(z^{(1)}), hidden error delta vector δ(1)\delta^{(1)}, hidden weight gradient matrix ∂L∂W(1)\frac{\partial L}{\partial W^{(1)}}, and hidden bias gradient ∂L∂b(1)\frac{\partial L}{\partial b^{(1)}}.
  • Part D (Single-Step Parameter Update): Compute the updated weight W11,new(1)W^{(1)}_{11, \text{new}} using step size η=0.50\eta = 0.50, as well as the full updated parameter matrices Wnew(1),bnew(1),Wnew(2),bnew(2)W^{(1)}_{\text{new}}, b^{(1)}_{\text{new}}, W^{(2)}_{\text{new}}, b^{(2)}_{\text{new}}.
  • Part E (Second Forward Pass & Mathematical Proof): Re-evaluate the complete forward pass with all updated parameters, compute updated loss L2L_2, and prove that L2<L1L_2 < L_1 (ΔL<0\Delta L < 0).
  • Part F (Venture Synthesis): In 2–3 sentences, explain why the gradient for Market Size (W12(1)→0.0440W^{(1)}_{12} \to 0.0440) is exactly twice as large as Team Experience (W11(1)→0.0220W^{(1)}_{11} \to 0.0220), and how backpropagation automatically assigns larger corrective penalties to larger input signals.
Reveal Solution

Part A: Forward Pass 1 and Initial Loss

  1. Layer 1 Pre-activations (z(1)z^{(1)}): z1(1)=0.5(1.0)+0.5(2.0)+0.0=0.5+1.0=1.5000z_1^{(1)} = 0.5(1.0) + 0.5(2.0) + 0.0 = 0.5 + 1.0 = \mathbf{1.5000} z2(1)=0.5(1.0)+0.5(2.0)+0.0=0.5+1.0=1.5000z_2^{(1)} = 0.5(1.0) + 0.5(2.0) + 0.0 = 0.5 + 1.0 = \mathbf{1.5000} z(1)=[1.50001.5000]z^{(1)} = \begin{bmatrix} 1.5000 \\ 1.5000 \end{bmatrix}

  2. Layer 1 Activations (a(1)a^{(1)}): a1(1)=a2(1)=σ(1.5)=11+e−1.5=11+0.223130≈0.817574a_1^{(1)} = a_2^{(1)} = \sigma(1.5) = \frac{1}{1 + e^{-1.5}} = \frac{1}{1 + 0.223130} \approx \mathbf{0.817574} a(1)=[0.8175740.817574]a^{(1)} = \begin{bmatrix} 0.817574 \\ 0.817574 \end{bmatrix}

  3. Layer 2 Pre-activation (z(2)z^{(2)}): z(2)=W(2)a(1)+b(2)=(0.5)(0.817574)+(0.5)(0.817574)+0.0=0.817574z^{(2)} = W^{(2)} a^{(1)} + b^{(2)} = (0.5)(0.817574) + (0.5)(0.817574) + 0.0 = \mathbf{0.817574}

  4. Layer 2 Output Prediction (a(2)a^{(2)}): a(2)=σ(0.817574)=11+e−0.817574=11+0.441500≈0.693721(69.37%)a^{(2)} = \sigma(0.817574) = \frac{1}{1 + e^{-0.817574}} = \frac{1}{1 + 0.441500} \approx \mathbf{0.693721} \quad (\mathbf{69.37\%})

  5. Initial Loss (L1L_1): L1=(a(2)−y)2=(0.693721−0.0)2≈0.481249L_1 = (a^{(2)} - y)^2 = (0.693721 - 0.0)^2 \approx \mathbf{0.481249}


Part B: Output Layer Error Attribution

  1. Loss Sensitivity: ∂L∂a(2)=2(a(2)−y)=2(0.693721−0.0)=1.387442\frac{\partial L}{\partial a^{(2)}} = 2(a^{(2)} - y) = 2(0.693721 - 0.0) = \mathbf{1.387442}

  2. Output Activation Slope: σ′(z(2))=a(2)(1−a(2))=0.693721×(1−0.693721)=0.693721×0.306279=0.212472\sigma'(z^{(2)}) = a^{(2)}(1 - a^{(2)}) = 0.693721 \times (1 - 0.693721) = 0.693721 \times 0.306279 = \mathbf{0.212472}

  3. Output Error Delta (δ(2)\delta^{(2)}): δ(2)=∂L∂a(2)⋅σ′(z(2))=1.387442×0.212472=0.294793\delta^{(2)} = \frac{\partial L}{\partial a^{(2)}} \cdot \sigma'(z^{(2)}) = 1.387442 \times 0.212472 = \mathbf{0.294793}

  4. Layer 2 Weight Gradients: ∂L∂W(2)=δ(2)(a(1))T=0.294793×[0.8175740.817574]=[0.2410150.241015]\frac{\partial L}{\partial W^{(2)}} = \delta^{(2)} (a^{(1)})^T = 0.294793 \times \begin{bmatrix} 0.817574 & 0.817574 \end{bmatrix} = \begin{bmatrix} \mathbf{0.241015} & \mathbf{0.241015} \end{bmatrix}

  5. Layer 2 Bias Gradient: ∂L∂b(2)=δ(2)=0.294793\frac{\partial L}{\partial b^{(2)}} = \delta^{(2)} = \mathbf{0.294793}


Part C: Hidden Layer Backpropagation

  1. Transposed Weight Projection: (W(2))Tδ(2)=[0.50.5](0.294793)=[0.1473960.147396](W^{(2)})^T \delta^{(2)} = \begin{bmatrix} 0.5 \\ 0.5 \end{bmatrix} (0.294793) = \begin{bmatrix} \mathbf{0.147396} \\ \mathbf{0.147396} \end{bmatrix}

  2. Hidden Activation Slopes: σ′(z1(1))=σ′(z2(1))=a1(1)(1−a1(1))=0.817574×(1−0.817574)=0.817574×0.182426=0.149146\sigma'(z_1^{(1)}) = \sigma'(z_2^{(1)}) = a_1^{(1)}(1 - a_1^{(1)}) = 0.817574 \times (1 - 0.817574) = 0.817574 \times 0.182426 = \mathbf{0.149146}

  3. Hidden Error Delta Vector (δ(1)\delta^{(1)}): δ(1)=[0.1473960.147396]⊙[0.1491460.149146]=[0.0219840.021984]\delta^{(1)} = \begin{bmatrix} 0.147396 \\ 0.147396 \end{bmatrix} \odot \begin{bmatrix} 0.149146 \\ 0.149146 \end{bmatrix} = \begin{bmatrix} \mathbf{0.021984} \\ \mathbf{0.021984} \end{bmatrix}

  4. Layer 1 Weight Gradients: ∂L∂W(1)=δ(1)xT=[0.0219840.021984][1.02.0]=[0.0219840.0439670.0219840.043967]\frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T = \begin{bmatrix} 0.021984 \\ 0.021984 \end{bmatrix} \begin{bmatrix} 1.0 & 2.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.021984} & \mathbf{0.043967} \\ \mathbf{0.021984} & \mathbf{0.043967} \end{bmatrix}

  5. Layer 1 Bias Gradient: ∂L∂b(1)=δ(1)=[0.0219840.021984]\frac{\partial L}{\partial b^{(1)}} = \delta^{(1)} = \begin{bmatrix} \mathbf{0.021984} \\ \mathbf{0.021984} \end{bmatrix}


Part D: Single-Step Parameter Updates with η=0.50\eta = 0.50

  1. Target Weight W11(1)W^{(1)}_{11} Update: W11,new(1)=W11,old(1)−η∂L∂W11(1)=0.5−0.5(0.021984)=0.5−0.010992=0.489008W^{(1)}_{11, \text{new}} = W^{(1)}_{11, \text{old}} - \eta \frac{\partial L}{\partial W^{(1)}_{11}} = 0.5 - 0.5(0.021984) = 0.5 - 0.010992 = \mathbf{0.489008}

  2. Full Parameter Updates: Wnew(1)=[0.50.50.50.5]−0.5[0.0219840.0439670.0219840.043967]=[0.4890080.4780160.4890080.478016]W^{(1)}_{\text{new}} = \begin{bmatrix} 0.5 & 0.5 \\ 0.5 & 0.5 \end{bmatrix} - 0.5 \begin{bmatrix} 0.021984 & 0.043967 \\ 0.021984 & 0.043967 \end{bmatrix} = \begin{bmatrix} \mathbf{0.489008} & \mathbf{0.478016} \\ \mathbf{0.489008} & \mathbf{0.478016} \end{bmatrix} bnew(1)=[0.00.0]−0.5[0.0219840.021984]=[−0.010992−0.010992]b^{(1)}_{\text{new}} = \begin{bmatrix} 0.0 \\ 0.0 \end{bmatrix} - 0.5 \begin{bmatrix} 0.021984 \\ 0.021984 \end{bmatrix} = \begin{bmatrix} \mathbf{-0.010992} \\ \mathbf{-0.010992} \end{bmatrix} Wnew(2)=[0.50.5]−0.5[0.2410150.241015]=[0.3794920.379492]W^{(2)}_{\text{new}} = \begin{bmatrix} 0.5 & 0.5 \end{bmatrix} - 0.5 \begin{bmatrix} 0.241015 & 0.241015 \end{bmatrix} = \begin{bmatrix} \mathbf{0.379492} & \mathbf{0.379492} \end{bmatrix} bnew(2)=0.0−0.5(0.294793)=−0.147396b^{(2)}_{\text{new}} = 0.0 - 0.5(0.294793) = \mathbf{-0.147396}


Part E: Second Forward Pass and Mathematical Loss Reduction Proof

  1. Updated Layer 1 Pass: z1(1)new=(0.489008)(1.0)+(0.478016)(2.0)−0.010992=0.489008+0.956033−0.010992=1.434049z_1^{(1)\text{new}} = (0.489008)(1.0) + (0.478016)(2.0) - 0.010992 = 0.489008 + 0.956033 - 0.010992 = \mathbf{1.434049} a1(1)new=a2(1)new=σ(1.434049)=11+e−1.434049≈0.807531a_1^{(1)\text{new}} = a_2^{(1)\text{new}} = \sigma(1.434049) = \frac{1}{1 + e^{-1.434049}} \approx \mathbf{0.807531}

  2. Updated Layer 2 Pass: z(2)new=(0.379492)(0.807531)+(0.379492)(0.807531)−0.147396=0.612904−0.147396=0.465508z^{(2)\text{new}} = (0.379492)(0.807531) + (0.379492)(0.807531) - 0.147396 = 0.612904 - 0.147396 = \mathbf{0.465508} a(2)new=σ(0.465508)=11+e−0.465508≈0.614320(61.43%)a^{(2)\text{new}} = \sigma(0.465508) = \frac{1}{1 + e^{-0.465508}} \approx \mathbf{0.614320} \quad (\mathbf{61.43\%})

  3. Updated Loss Evaluation (L2L_2): L2=(a(2)new−y)2=(0.614320−0.0)2≈0.377389L_2 = (a^{(2)\text{new}} - y)^2 = (0.614320 - 0.0)^2 \approx \mathbf{0.377389}

  4. Proof of Strict Loss Reduction: ΔL=L2−L1=0.377389−0.481249=−0.103860<0\Delta L = L_2 - L_1 = 0.377389 - 0.481249 = \mathbf{-0.103860} < 0 Relative Error Reduction=0.1038600.481249×100%=21.58%\text{Relative Error Reduction} = \frac{0.103860}{0.481249} \times 100\% = \mathbf{21.58\%}

The parameter update step reduced prediction loss by 21.58%21.58\%, pulling the network's failure probability estimate from 69.37%69.37\% down to 61.43%61.43\%.


Part F: Venture Synthesis

The gradient for Market Size (W12(1)→0.0440W^{(1)}_{12} \to 0.0440) is exactly twice as large as Team Experience (W11(1)→0.0220W^{(1)}_{11} \to 0.0220) because the input measurement for Market Size was twice as large (x2=2.0x_2 = 2.0 vs x1=1.0x_1 = 1.0).

Because the parameter gradient is given by ∂L∂W(1)=δ(1)xT\frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T, features with larger input magnitudes amplify the error signal and receive proportionately larger gradient penalties during backpropagation.

Previous
Toy Parameter Update Step on Paper