Output Loss Derivatives and Gradients hero
Lesson 4Prediction Error & Loss Functions

Output Loss Derivatives and Gradients

Compute the derivative of the loss function with respect to predicted outputs, establishing the first mathematical link in error attribution.

Knowing our current loss altitude (L=0.9025L = 0.9025) tells us how wrong the model was, but it does not tell the network which direction to adjust its prediction to reduce the error.

To find that direction, we compute the rate of change of the loss with respect to the prediction: the derivative ∂L∂y^\frac{\partial L}{\partial \hat{y}}.


The Rate-of-Change Question

We ask a fundamental calculus question:

"If we nudge our prediction y^ upward by a tiny amount Δy^, how much does loss L change?"\text{\textit{"If we nudge our prediction $\hat{y}$ upward by a tiny amount $\Delta \hat{y}$, how much does loss $L$ change?"}}

Mathematically, this sensitivity is given by the instantaneous derivative:

∂L∂y^=lim⁡Δy^→0L(y^+Δy^)−L(y^)Δy^\frac{\partial L}{\partial \hat{y}} = \lim_{\Delta \hat{y} \to 0} \frac{L(\hat{y} + \Delta \hat{y}) - L(\hat{y})}{\Delta \hat{y}}

Step-by-Step Analytical Calculus Derivation

Let's derive ∂L∂y^\frac{\partial L}{\partial \hat{y}} from first principles using the Chain Rule.

We express standard squared loss as:

L=(y^−y)2L = (\hat{y} - y)^2

Let the intermediate variable be u=y^−yu = \hat{y} - y, so that L=u2L = u^2.

Applying the Chain Rule:

∂L∂y^=dLdu⋅∂u∂y^\frac{\partial L}{\partial \hat{y}} = \frac{dL}{du} \cdot \frac{\partial u}{\partial \hat{y}}
  1. Outer Derivative: dLdu=ddu[u2]=2u=2(y^−y)\frac{dL}{du} = \frac{d}{du}[u^2] = 2u = 2(\hat{y} - y)

  2. Inner Derivative: ∂u∂y^=∂∂y^[y^−y]=1−0=1\frac{\partial u}{\partial \hat{y}} = \frac{\partial}{\partial \hat{y}}[\hat{y} - y] = 1 - 0 = 1

Multiplying outer and inner derivatives yields the standard output loss derivative:

∂L∂y^=2(y^−y)⋅1=2(y^−y)\frac{\partial L}{\partial \hat{y}} = 2(\hat{y} - y) \cdot 1 = \mathbf{2(\hat{y} - y)}

Derivation Starting from L=(y−y^)2L = (y - \hat{y})^2:

If we define loss as L=(y−y^)2L = (y - \hat{y})^2 and let v=y−y^v = y - \hat{y}:

  • Outer derivative: dLdv=2v=2(y−y^)\frac{dL}{dv} = 2v = 2(y - \hat{y})
  • Inner derivative: ∂v∂y^=∂∂y^[y−y^]=0−1=−1\frac{\partial v}{\partial \hat{y}} = \frac{\partial}{\partial \hat{y}}[y - \hat{y}] = 0 - 1 = -1

Multiplying together:

∂L∂y^=2(y−y^)⋅(−1)=−2(y−y^)=2(y^−y)\frac{\partial L}{\partial \hat{y}} = 2(y - \hat{y}) \cdot (-1) = -2(y - \hat{y}) = \mathbf{2(\hat{y} - y)}

Both starting formulations produce the exact same derivative: ∂L∂y^=2(y^−y)\frac{\partial L}{\partial \hat{y}} = 2(\hat{y} - y).

Derivation for Scaled Loss (Lscaled=12(y^−y)2L_{\text{scaled}} = \frac{1}{2}(\hat{y} - y)^2):

Under the 12\frac{1}{2} calculus scaling convention:

∂Lscaled∂y^=∂∂y^[12(y^−y)2]=12⋅2(y^−y)=y^−y\frac{\partial L_{\text{scaled}}}{\partial \hat{y}} = \frac{\partial}{\partial \hat{y}}\left[ \frac{1}{2}(\hat{y} - y)^2 \right] = \frac{1}{2} \cdot 2(\hat{y} - y) = \mathbf{\hat{y} - y}

The factor of 22 cancels, leaving the derivative equal to the raw discrepancy y^−y\hat{y} - y.


Sign and Magnitude Interpretation

The sign and magnitude of ∂L∂y^\frac{\partial L}{\partial \hat{y}} provide an unambiguous navigation instruction for optimization:

Prediction RegimeConditionExample (y^,y\hat{y}, y)Derivative ∂L∂y^\frac{\partial L}{\partial \hat{y}}Slope DirectionCorrective Action to Reduce Loss
Overestimatey^>y\hat{y} > yy^=0.95,y=0.0\hat{y} = 0.95, y = 0.0+1.90+1.90 (Positive)Uphill (Loss increases as y^↑\hat{y} \uparrow)Decrease y^\hat{y} (Nudge left toward 00)
Underestimatey^<y\hat{y} < yy^=0.10,y=1.0\hat{y} = 0.10, y = 1.0−1.80-1.80 (Negative)Downhill (Loss decreases as y^↑\hat{y} \uparrow)Increase y^\hat{y} (Nudge right toward 11)
Perfect Matchy^=y\hat{y} = yy^=1.00,y=1.0\hat{y} = 1.00, y = 1.00.000.00 (Zero)Flat Basin (Stationary minimum)No change needed
Slope dL/dy_hat > 0 (Overestimate)   ──► Step LEFT  (Decrease y_hat)
Slope dL/dy_hat < 0 (Underestimate)  ──► Step RIGHT (Increase y_hat)
Slope dL/dy_hat = 0 (Target Hit)     ──► STOP       (Optimal Minimum)

In gradient descent, we always update variables in the opposite direction of the derivative (the negative gradient −∂L∂y^-\frac{\partial L}{\partial \hat{y}}):

Δy^∝−∂L∂y^=−2(y^−y)=2(y−y^)\Delta \hat{y} \propto -\frac{\partial L}{\partial \hat{y}} = -2(\hat{y} - y) = 2(y - \hat{y})
  • For 'Die Hard in Space' (y^=0.95,y=0\hat{y} = 0.95, y = 0): ∂L∂y^=+1.90  ⟹  −∂L∂y^=−1.90\frac{\partial L}{\partial \hat{y}} = +1.90 \implies -\frac{\partial L}{\partial \hat{y}} = \mathbf{-1.90}. The negative update direction instructs the network to pull its output activation downward toward zero.

The Upstream Error Attribution Link

In a neural network, we do not directly dial the output prediction y^\hat{y} by hand. The prediction is generated by output pre-activation z(2)z^{(2)}, which is computed from hidden layer activations a(1)a^{(1)}, which are computed from weight matrices W(1),W(2)W^{(1)}, W^{(2)} and biases b(1),b(2)b^{(1)}, b^{(2)}.

How does the error penalty at the output reach the millions of weights deep inside the network?

The output loss derivative ∂L∂y^\frac{\partial L}{\partial \hat{y}} is the starting spark of backpropagation.

[ Loss L ] ──► [ Output y_hat ] ──► [ Pre-activation z ] ──► [ Weights w ]
   dL/dL              dL/dy_hat            dy_hat/dz                dz/dw
   (= 1.0)            (Topic 1)            (Topic 2)              (Topic 3)

Using the multivariable chain rule, the sensitivity of the loss to any internal weight ww is factored into a sequence of chained derivatives:

∂L∂w=∂L∂y^⏟Topic 1: Loss Sensitivity⋅∂y^∂z⏟Topic 2: Activation Derivative⋅∂z∂w⏟Topic 3: Linear Weight Derivative\frac{\partial L}{\partial w} = \underbrace{\frac{\partial L}{\partial \hat{y}}}_{\text{Topic 1: Loss Sensitivity}} \cdot \underbrace{\frac{\partial \hat{y}}{\partial z}}_{\text{Topic 2: Activation Derivative}} \cdot \underbrace{\frac{\partial z}{\partial w}}_{\text{Topic 3: Linear Weight Derivative}}

By calculating ∂L∂y^=+1.90\frac{\partial L}{\partial \hat{y}} = +1.90 today, we have established the very first link in the backward error chain. In the remaining topics of Module 2, we will propagate this numerical signal backward through the activation functions and matrix transformations to compute exact gradient updates for every weight in the network.


Previous
Parabolic Loss Surfaces and Curvature