Toy Parameter Update Step on Paper hero
Lesson 4The Chain Rule and Backpropagation

Toy Parameter Update Step on Paper

Perform a single conceptual parameter update step by scaling computed gradients with a step factor to demonstrate how weights adjust on paper.

With all parameter gradients computed, we can now perform a single conceptual parameter update step to observe how the network learns on paper.


The Parameter Update Rule

In gradient descent, every parameter in the network is adjusted by taking a small step in the direction opposite to its computed gradient:

Wnew=Wold−η∂L∂WorW←W−η∇WLW_{\text{new}} = W_{\text{old}} - \eta \frac{\partial L}{\partial W} \quad \text{or} \quad W \leftarrow W - \eta \nabla_W L bnew=bold−η∂L∂borb←b−η∇bLb_{\text{new}} = b_{\text{old}} - \eta \frac{\partial L}{\partial b} \quad \text{or} \quad b \leftarrow b - \eta \nabla_b L

For an individual scalar weight ww:

wnew=wold−η∂L∂ww_{\text{new}} = w_{\text{old}} - \eta \frac{\partial L}{\partial w}

For layer ll in a multi-layer neural network:

Wnew(l)=Wold(l)−η∂L∂W(l)W^{(l)}_{\text{new}} = W^{(l)}_{\text{old}} - \eta \frac{\partial L}{\partial W^{(l)}} bnew(l)=bold(l)−η∂L∂b(l)b^{(l)}_{\text{new}} = b^{(l)}_{\text{old}} - \eta \frac{\partial L}{\partial b^{(l)}}

Where η>0\eta > 0 is the Step Size Factor (commonly termed the learning rate).

Why Do We Need a Step Size Factor η\eta?

  • The gradient ∂L∂W\frac{\partial L}{\partial W} represents the instantaneous slope at the current parameter coordinate.
  • If we took a full unscaled step (η=1.0\eta = 1.0), the update could wildly overshoot the minimum of the error bowl.
  • Scaling the gradient by a calibrated step size (such as η=0.5\eta = 0.5 or η=0.1\eta = 0.1) ensures controlled, stable progress down the loss surface.

Step-by-Step Hand Calculation on 'Die Hard in Space'

Let us apply a single gradient descent update step with step size η=0.5\mathbf{\eta = 0.5} to all parameters in our 2-layer network.

1. Layer 2 Parameter Updates

  • Weight Update: Wnew(2)=[1.51.0]−0.5[0.1387600.000000]=[1.4306201.000000]W^{(2)}_{\text{new}} = \begin{bmatrix} 1.5 & 1.0 \end{bmatrix} - 0.5 \begin{bmatrix} 0.138760 & 0.000000 \end{bmatrix} = \begin{bmatrix} \mathbf{1.430620} & \mathbf{1.000000} \end{bmatrix}
  • Bias Update: bnew(2)=−2.0−0.5(0.034690)=−2.017345b^{(2)}_{\text{new}} = -2.0 - 0.5(0.034690) = \mathbf{-2.017345}

2. Layer 1 Parameter Updates

  • Weight Matrix Update: Wnew(1)=[3.0−2.00.03.0−2.03.02.0−1.0]−0.5[0.0520350.0000000.0260180.0520350.0000000.0000000.0000000.000000]W^{(1)}_{\text{new}} = \begin{bmatrix} 3.0 & -2.0 & 0.0 & 3.0 \\ -2.0 & 3.0 & 2.0 & -1.0 \end{bmatrix} - 0.5 \begin{bmatrix} 0.052035 & 0.000000 & 0.026018 & 0.052035 \\ 0.000000 & 0.000000 & 0.000000 & 0.000000 \end{bmatrix} Wnew(1)=[2.973982−2.000000−0.0130092.973982−2.0000003.0000002.000000−1.000000]W^{(1)}_{\text{new}} = \begin{bmatrix} \mathbf{2.973982} & \mathbf{-2.000000} & \mathbf{-0.013009} & \mathbf{2.973982} \\ \mathbf{-2.000000} & \mathbf{3.000000} & \mathbf{2.000000} & \mathbf{-1.000000} \end{bmatrix}
  • Bias Vector Update: bnew(1)=[−2.0−1.5]−0.5[0.0520350.000000]=[−2.026018−1.500000]b^{(1)}_{\text{new}} = \begin{bmatrix} -2.0 \\ -1.5 \end{bmatrix} - 0.5 \begin{bmatrix} 0.052035 \\ 0.000000 \end{bmatrix} = \begin{bmatrix} \mathbf{-2.026018} \\ \mathbf{-1.500000} \end{bmatrix}

Second Forward Pass: Proving Mathematical Loss Reduction

To prove that this single backpropagation update step genuinely improved the model, let us re-run the complete forward pass using our updated parameters.

Step 1: Updated Hidden Layer Forward Pass

Compute znew(1)=Wnew(1)x+bnew(1)z^{(1)}_{\text{new}} = W^{(1)}_{\text{new}} x + b^{(1)}_{\text{new}}:

z1,new(1)=(2.973982)(1.0)+(−2.0)(0.0)+(−0.013009)(0.5)+(2.973982)(1.0)−2.026018z_{1, \text{new}}^{(1)} = (2.973982)(1.0) + (-2.0)(0.0) + (-0.013009)(0.5) + (2.973982)(1.0) - 2.026018 z1,new(1)=2.973982−0.006505+2.973982−2.026018=+3.915441z_{1, \text{new}}^{(1)} = 2.973982 - 0.006505 + 2.973982 - 2.026018 = \mathbf{+3.915441} z2,new(1)=−3.500000z_{2, \text{new}}^{(1)} = -3.500000

Apply ReLU activation:

anew(1)=[max⁡(0,3.915441)max⁡(0,−3.5)]=[3.9154410.000000]a^{(1)}_{\text{new}} = \begin{bmatrix} \max(0, 3.915441) \\ \max(0, -3.5) \end{bmatrix} = \begin{bmatrix} \mathbf{3.915441} \\ \mathbf{0.000000} \end{bmatrix}

Step 2: Updated Output Layer Forward Pass

Compute znew(2)=Wnew(2)anew(1)+bnew(2)z^{(2)}_{\text{new}} = W^{(2)}_{\text{new}} a^{(1)}_{\text{new}} + b^{(2)}_{\text{new}}:

znew(2)=(1.430620)(3.915441)+(1.0)(0.0)−2.017345=5.601510−2.017345=+3.584165z^{(2)}_{\text{new}} = (1.430620)(3.915441) + (1.0)(0.0) - 2.017345 = 5.601510 - 2.017345 = \mathbf{+3.584165}

Apply Sigmoid activation:

anew(2)=σ(3.584165)=11+e−3.584165=11+0.027760=0.972990(97.30%)a^{(2)}_{\text{new}} = \sigma(3.584165) = \frac{1}{1 + e^{-3.584165}} = \frac{1}{1 + 0.027760} = \mathbf{0.972990} \quad (\mathbf{97.30\%})

Step 3: Evaluate Updated Loss Penalty

Compute the new Mean Squared Error loss:

Lnew=(anew(2)−y)2=(0.972990−0.0)2=0.946710L_{\text{new}} = (a^{(2)}_{\text{new}} - y)^2 = (0.972990 - 0.0)^2 = \mathbf{0.946710}

Mathematical Proof of Error Reduction

Compare the initial loss against the updated loss:

ΔL=Lnew−Lold=0.946710−0.964351=−0.017641(Strict Loss Reduction)\Delta L = L_{\text{new}} - L_{\text{old}} = 0.946710 - 0.964351 = \mathbf{-0.017641} \quad (\text{Strict Loss Reduction}) Relative Error Eliminated=0.0176410.964351×100%=1.83%\text{Relative Error Eliminated} = \frac{0.017641}{0.964351} \times 100\% = \mathbf{1.83\%}

The single backpropagation update step successfully nudged the prediction downward from 98.20%98.20\% toward the target of 0.0%0.0\%, strictly reducing prediction loss on paper.


Strict Scope Containment

We deliberately conclude our mathematical derivations at this single manual calculation.

In production engineering, training neural networks involves repeating this forward-backward cycle across thousands of batches. However, to maintain pedagogical clarity, the following advanced systems belong to subsequent courses:

  • ❌ No Training Loops or Epochs: No automated convergence loops or batch iterators.
  • ❌ No Complex Optimizers: No Momentum, Adam, or RMSProp algorithms.
  • ❌ No Autograd Frameworks: No PyTorch, TensorFlow, or JAX code abstractions.

Mastering the exact mechanics on paper guarantees that when you later encounter automated engines, you understand the exact linear algebra executing beneath every line of code.


Previous
Hidden Layer Error Backpropagation