With all parameter gradients computed, we can now perform a single conceptual parameter update step to observe how the network learns on paper.
The Parameter Update Rule
In gradient descent, every parameter in the network is adjusted by taking a small step in the direction opposite to its computed gradient:
Wnew=Wold−η∂W∂LorW←W−η∇WL
bnew=bold−η∂b∂Lorb←b−η∇bL
For an individual scalar weight w:
wnew=wold−η∂w∂L
For layer l in a multi-layer neural network:
Wnew(l)=Wold(l)−η∂W(l)∂L
bnew(l)=bold(l)−η∂b(l)∂L
Where η>0 is the Step Size Factor (commonly termed the learning rate).
Why Do We Need a Step Size Factor η?
- The gradient ∂W∂L represents the instantaneous slope at the current parameter coordinate.
- If we took a full unscaled step (η=1.0), the update could wildly overshoot the minimum of the error bowl.
- Scaling the gradient by a calibrated step size (such as η=0.5 or η=0.1) ensures controlled, stable progress down the loss surface.
Step-by-Step Hand Calculation on 'Die Hard in Space'
Let us apply a single gradient descent update step with step size η=0.5 to all parameters in our 2-layer network.
1. Layer 2 Parameter Updates
- Weight Update:
Wnew(2)=[1.51.0]−0.5[0.1387600.000000]=[1.4306201.000000]
- Bias Update:
bnew(2)=−2.0−0.5(0.034690)=−2.017345
2. Layer 1 Parameter Updates
- Weight Matrix Update:
Wnew(1)=[3.0−2.0−2.03.00.02.03.0−1.0]−0.5[0.0520350.0000000.0000000.0000000.0260180.0000000.0520350.000000]
Wnew(1)=[2.973982−2.000000−2.0000003.000000−0.0130092.0000002.973982−1.000000]
- Bias Vector Update:
bnew(1)=[−2.0−1.5]−0.5[0.0520350.000000]=[−2.026018−1.500000]
Second Forward Pass: Proving Mathematical Loss Reduction
To prove that this single backpropagation update step genuinely improved the model, let us re-run the complete forward pass using our updated parameters.
Step 1: Updated Hidden Layer Forward Pass
Compute znew(1)=Wnew(1)x+bnew(1):
z1,new(1)=(2.973982)(1.0)+(−2.0)(0.0)+(−0.013009)(0.5)+(2.973982)(1.0)−2.026018
z1,new(1)=2.973982−0.006505+2.973982−2.026018=+3.915441
z2,new(1)=−3.500000
Apply ReLU activation:
anew(1)=[max(0,3.915441)max(0,−3.5)]=[3.9154410.000000]
Step 2: Updated Output Layer Forward Pass
Compute znew(2)=Wnew(2)anew(1)+bnew(2):
znew(2)=(1.430620)(3.915441)+(1.0)(0.0)−2.017345=5.601510−2.017345=+3.584165
Apply Sigmoid activation:
anew(2)=σ(3.584165)=1+e−3.5841651=1+0.0277601=0.972990(97.30%)
Step 3: Evaluate Updated Loss Penalty
Compute the new Mean Squared Error loss:
Lnew=(anew(2)−y)2=(0.972990−0.0)2=0.946710
Mathematical Proof of Error Reduction
Compare the initial loss against the updated loss:
ΔL=Lnew−Lold=0.946710−0.964351=−0.017641(Strict Loss Reduction)
Relative Error Eliminated=0.9643510.017641×100%=1.83%
The single backpropagation update step successfully nudged the prediction downward from 98.20% toward the target of 0.0%, strictly reducing prediction loss on paper.
Strict Scope Containment
We deliberately conclude our mathematical derivations at this single manual calculation.
In production engineering, training neural networks involves repeating this forward-backward cycle across thousands of batches. However, to maintain pedagogical clarity, the following advanced systems belong to subsequent courses:
- ❌ No Training Loops or Epochs: No automated convergence loops or batch iterators.
- ❌ No Complex Optimizers: No Momentum, Adam, or RMSProp algorithms.
- ❌ No Autograd Frameworks: No PyTorch, TensorFlow, or JAX code abstractions.
Mastering the exact mechanics on paper guarantees that when you later encounter automated engines, you understand the exact linear algebra executing beneath every line of code.