Compute the derivative of the loss function with respect to predicted outputs, establishing the first mathematical link in error attribution.
Knowing our current loss altitude (L=0.9025) tells us how wrong the model was, but it does not tell the network which direction to adjust its prediction to reduce the error.
To find that direction, we compute the rate of change of the loss with respect to the prediction: the derivative ∂y^∂L.
The Rate-of-Change Question
We ask a fundamental calculus question:
"If we nudge our prediction y^ upward by a tiny amount Δy^, how much does loss L change?"
Mathematically, this sensitivity is given by the instantaneous derivative:
∂y^∂L=Δy^→0limΔy^L(y^+Δy^)−L(y^)
Step-by-Step Analytical Calculus Derivation
Let's derive ∂y^∂L from first principles using the Chain Rule.
We express standard squared loss as:
L=(y^−y)2
Let the intermediate variable be u=y^−y, so that L=u2.
Applying the Chain Rule:
∂y^∂L=dudL⋅∂y^∂u
Outer Derivative:dudL=dud[u2]=2u=2(y^−y)
Inner Derivative:∂y^∂u=∂y^∂[y^−y]=1−0=1
Multiplying outer and inner derivatives yields the standard output loss derivative:
∂y^∂L=2(y^−y)⋅1=2(y^−y)
Derivation Starting from L=(y−y^)2:
If we define loss as L=(y−y^)2 and let v=y−y^:
Outer derivative: dvdL=2v=2(y−y^)
Inner derivative: ∂y^∂v=∂y^∂[y−y^]=0−1=−1
Multiplying together:
∂y^∂L=2(y−y^)⋅(−1)=−2(y−y^)=2(y^−y)
Both starting formulations produce the exact same derivative: ∂y^∂L=2(y^−y).
Derivation for Scaled Loss (Lscaled=21(y^−y)2):
In gradient descent, we always update variables in the opposite direction of the derivative (the negative gradient −∂y^∂L):
Δy^∝−∂y^∂L=−2(y^−y)=2(y−y^)
For 'Die Hard in Space' (y^=0.95,y=0): ∂y^∂L=+1.90⟹−∂y^∂L=−1.90. The negative update direction instructs the network to pull its output activation downward toward zero.
The Upstream Error Attribution Link
In a neural network, we do not directly dial the output prediction y^ by hand. The prediction is generated by output pre-activation z(2), which is computed from hidden layer activations a(1), which are computed from weight matrices W(1),W(2) and biases b(1),b(2).
How does the error penalty at the output reach the millions of weights deep inside the network?
The output loss derivative ∂y^∂L is the starting spark of backpropagation.
[ Loss L ] ──► [ Output y_hat ] ──► [ Pre-activation z ] ──► [ Weights w ]
dL/dL dL/dy_hat dy_hat/dz dz/dw
(= 1.0) (Topic 1) (Topic 2) (Topic 3)
Using the multivariable chain rule, the sensitivity of the loss to any internal weight w is factored into a sequence of chained derivatives:
∂w∂L=Topic 1: Loss Sensitivity∂y^∂L⋅Topic 2: Activation Derivative∂z∂y^⋅Topic 3: Linear Weight Derivative∂w∂z
By calculating ∂y^∂L=+1.90 today, we have established the very first link in the backward error chain. In the remaining topics of Module 2, we will propagate this numerical signal backward through the activation functions and matrix transformations to compute exact gradient updates for every weight in the network.