Output Layer Error Attribution Math hero
Lesson 2The Chain Rule and Backpropagation

Output Layer Error Attribution Math

Decompose output layer gradients into loss sensitivity, activation slope, and input terms to compute exact credit attribution for output weights.

Now let us apply the chain rule to the output layer of our neural network.

Consider a 2-layer network. The output layer receives a vector of hidden activations a(1)∈Rd1×1a^{(1)} \in \mathbb{R}^{d_1 \times 1}, combines them with output weight matrix W(2)∈R1×d1W^{(2)} \in \mathbb{R}^{1 \times d_1} and output bias b(2)∈R1×1b^{(2)} \in \mathbb{R}^{1 \times 1}, passes the result through an activation function σ\sigma, and evaluates prediction error against ground truth yy:

Hidden Activations a^(1) ──► [ Affine Sum: z^(2) = W^(2) a^(1) + b^(2) ] ──► [ Activation: a^(2) = σ(z^(2)) ] ──► [ Loss: L = (a^(2) - y)^2 ]

To update the output weights, we must compute the parameter gradient matrix ∂L∂W(2)\frac{\partial L}{\partial W^{(2)}}.


The 3-Factor Gradient Decomposition

By applying the chain rule, the gradient of the loss with respect to output weight Wij(2)W^{(2)}_{ij} decomposes into three chained derivative factors:

∂L∂Wij(2)=∂L∂a(2)⏟Factor 1: Loss Sensitivity⋅∂a(2)∂z(2)⏟Factor 2: Activation Slope⋅∂z(2)∂Wij(2)⏟Factor 3: Affine Input\frac{\partial L}{\partial W^{(2)}_{ij}} = \underbrace{\frac{\partial L}{\partial a^{(2)}}}_{\text{Factor 1: Loss Sensitivity}} \cdot \underbrace{\frac{\partial a^{(2)}}{\partial z^{(2)}}}_{\text{Factor 2: Activation Slope}} \cdot \underbrace{\frac{\partial z^{(2)}}{\partial W^{(2)}_{ij}}}_{\text{Factor 3: Affine Input}}

Let us derive each of these three factors individually:

Factor 1: Loss Sensitivity (∂L∂a(2)\frac{\partial L}{\partial a^{(2)}})

For standard Mean Squared Error loss L=(a(2)−y)2L = (a^{(2)} - y)^2:

∂L∂a(2)=dda(2)[(a(2)−y)2]=2(a(2)−y)\frac{\partial L}{\partial a^{(2)}} = \frac{d}{da^{(2)}} \left[ (a^{(2)} - y)^2 \right] = 2(a^{(2)} - y)

This term measures the raw prediction error direction and magnitude. If a(2)>ya^{(2)} > y, the loss sensitivity is positive, signaling that the output was too high.

Factor 2: Activation Function Slope (σ′(z(2))\sigma'(z^{(2)}))

For the Sigmoid squashing function a(2)=σ(z(2))=11+e−z(2)a^{(2)} = \sigma(z^{(2)}) = \frac{1}{1 + e^{-z^{(2)}}}, we compute its derivative using the power and chain rules:

σ(z)=(1+e−z)−1\sigma(z) = (1 + e^{-z})^{-1} ddzσ(z)=−(1+e−z)−2⋅(−e−z)=e−z(1+e−z)2=(11+e−z)(e−z1+e−z)\frac{d}{dz}\sigma(z) = -(1 + e^{-z})^{-2} \cdot (-e^{-z}) = \frac{e^{-z}}{(1 + e^{-z})^2} = \left(\frac{1}{1 + e^{-z}}\right) \left(\frac{e^{-z}}{1 + e^{-z}}\right)

Since e−z1+e−z=1+e−z−11+e−z=1−σ(z)\frac{e^{-z}}{1 + e^{-z}} = \frac{1 + e^{-z} - 1}{1 + e^{-z}} = 1 - \sigma(z), we obtain the Sigmoid derivative identity:

σ′(z)=σ(z)(1−σ(z))=a(2)(1−a(2))\sigma'(z) = \sigma(z)(1 - \sigma(z)) = a^{(2)}(1 - a^{(2)})

This factor acts as a sensitivity gate. When the neuron is saturated near 00 or 11, the slope σ′(z)\sigma'(z) approaches zero, dampening the error signal.

Factor 3: Affine Input Sensitivity (∂z(2)∂W(2)\frac{\partial z^{(2)}}{\partial W^{(2)}})

The linear pre-activation is z(2)=∑jW1j(2)aj(1)+b(2)z^{(2)} = \sum_j W^{(2)}_{1j} a^{(1)}_j + b^{(2)}. Differentiating with respect to weight W1j(2)W^{(2)}_{1j} holding all other terms constant:

∂z(2)∂W1j(2)=aj(1)\frac{\partial z^{(2)}}{\partial W^{(2)}_{1j}} = a^{(1)}_j

And differentiating with respect to the bias b(2)b^{(2)}:

∂z(2)∂b(2)=1\frac{\partial z^{(2)}}{\partial b^{(2)}} = 1

The Output Layer Error Delta (δ(2)\delta^{(2)})

To keep calculations organized and avoid redundant multiplications, deep learning packages Factors 1 and 2 into a single intermediate term called the Error Delta (denoted δ\delta):

δ(2)≡∂L∂z(2)=∂L∂a(2)⋅∂a(2)∂z(2)=2(a(2)−y)⋅σ(z(2))(1−σ(z(2)))\delta^{(2)} \equiv \frac{\partial L}{\partial z^{(2)}} = \frac{\partial L}{\partial a^{(2)}} \cdot \frac{\partial a^{(2)}}{\partial z^{(2)}} = 2(a^{(2)} - y) \cdot \sigma(z^{(2)})(1 - \sigma(z^{(2)}))

In vector notation (using the Hadamard element-wise product ⊙\odot for multi-output layers):

δ(2)=∂L∂a(2)⊙σ′(z(2))\delta^{(2)} = \frac{\partial L}{\partial a^{(2)}} \odot \sigma'(z^{(2)})

The error delta δ(2)\delta^{(2)} represents the exact rate of change of the loss with respect to the pre-activation z(2)z^{(2)}. It quantifies how much total blame arrives at the output layer's linear sum.


Output Parameter Gradients

Once δ(2)\delta^{(2)} is computed, obtaining the parameter gradients requires only multiplying by the incoming inputs:

∂L∂W(2)=δ(2)(a(1))T\frac{\partial L}{\partial W^{(2)}} = \delta^{(2)} (a^{(1)})^T ∂L∂b(2)=δ(2)\frac{\partial L}{\partial b^{(2)}} = \delta^{(2)}

Dimensional Consistency Audit Table

To ensure all operations are valid, let us check matrix shapes for our 4→2→14 \to 2 \to 1 network:

VariableMathematical MeaningLocal Shape (4→2→14 \to 2 \to 1)General Shape (d0→d1→d2d_0 \to d_1 \to d_2)
a(1)a^{(1)}Hidden Layer Activation VectorR2×1\mathbb{R}^{2 \times 1}Rd1×1\mathbb{R}^{d_1 \times 1}
(a(1))T(a^{(1)})^TTransposed Hidden ActivationR1×2\mathbb{R}^{1 \times 2}R1×d1\mathbb{R}^{1 \times d_1}
z(2)z^{(2)}Output Pre-activationR1×1\mathbb{R}^{1 \times 1}Rd2×1\mathbb{R}^{d_2 \times 1}
a(2)a^{(2)}Output Activation PredictionR1×1\mathbb{R}^{1 \times 1}Rd2×1\mathbb{R}^{d_2 \times 1}
∂L∂a(2)\frac{\partial L}{\partial a^{(2)}}Loss Sensitivity VectorR1×1\mathbb{R}^{1 \times 1}Rd2×1\mathbb{R}^{d_2 \times 1}
σ′(z(2))\sigma'(z^{(2)})Activation Slope VectorR1×1\mathbb{R}^{1 \times 1}Rd2×1\mathbb{R}^{d_2 \times 1}
δ(2)\delta^{(2)}Output Error DeltaR1×1\mathbb{R}^{1 \times 1}Rd2×1\mathbb{R}^{d_2 \times 1}
∂L∂W(2)\frac{\partial L}{\partial W^{(2)}}Output Weight Gradient MatrixR1×1×R1×2=R1×2\mathbb{R}^{1 \times 1} \times \mathbb{R}^{1 \times 2} = \mathbf{\mathbb{R}^{1 \times 2}}Rd2×1×R1×d1=Rd2×d1\mathbb{R}^{d_2 \times 1} \times \mathbb{R}^{1 \times d_1} = \mathbf{\mathbb{R}^{d_2 \times d_1}}
∂L∂b(2)\frac{\partial L}{\partial b^{(2)}}Output Bias Gradient VectorR1×1\mathbf{\mathbb{R}^{1 \times 1}}Rd2×1\mathbf{\mathbb{R}^{d_2 \times 1}}

The gradient matrix ∂L∂W(2)∈R1×2\frac{\partial L}{\partial W^{(2)}} \in \mathbb{R}^{1 \times 2} matches the exact dimensions of weight matrix W(2)∈R1×2W^{(2)} \in \mathbb{R}^{1 \times 2}.


Step-by-Step Numerical Walkthrough: 'Die Hard in Space'

Let us compute the exact output layer gradients for our running movie prediction network on 'Die Hard in Space'.

  • Ground-Truth Target: y=0y = 0 (Box Office Flop)
  • Forward Pass State from Module 1 Topic 4:
    • Hidden activation vector: a(1)=[4.00.0]a^{(1)} = \begin{bmatrix} 4.0 \\ 0.0 \end{bmatrix} (Neuron 1: Summer Popcorn Flick, Neuron 2: Rom-Com)
    • Layer 2 weights: W(2)=[1.51.0]W^{(2)} = \begin{bmatrix} 1.5 & 1.0 \end{bmatrix}, bias b(2)=−2.0b^{(2)} = -2.0
    • Output pre-activation: z(2)=(1.5)(4.0)+(1.0)(0.0)−2.0=+4.0z^{(2)} = (1.5)(4.0) + (1.0)(0.0) - 2.0 = +4.0
    • Output prediction: a(2)=σ(4.0)=11+e−4.0≈0.982014a^{(2)} = \sigma(4.0) = \frac{1}{1 + e^{-4.0}} \approx \mathbf{0.982014} (98.20%98.20\% hit probability)
    • Mean Squared Error Loss: L=(0.982014−0)2≈0.964351L = (0.982014 - 0)^2 \approx \mathbf{0.964351}

Step 1: Compute Loss Sensitivity

∂L∂a(2)=2(a(2)−y)=2(0.982014−0.0)=+1.964028\frac{\partial L}{\partial a^{(2)}} = 2(a^{(2)} - y) = 2(0.982014 - 0.0) = \mathbf{+1.964028}

Step 2: Compute Activation Slope

σ′(z(2))=a(2)(1−a(2))=0.982014×(1−0.982014)=0.982014×0.017986=0.017663\sigma'(z^{(2)}) = a^{(2)}(1 - a^{(2)}) = 0.982014 \times (1 - 0.982014) = 0.982014 \times 0.017986 = \mathbf{0.017663}

Step 3: Compute Output Error Delta δ(2)\delta^{(2)}

δ(2)=∂L∂a(2)⋅σ′(z(2))=(+1.964028)×(0.017663)=+0.034690\delta^{(2)} = \frac{\partial L}{\partial a^{(2)}} \cdot \sigma'(z^{(2)}) = (+1.964028) \times (0.017663) = \mathbf{+0.034690}

Step 4: Compute Output Parameter Gradients

Compute the weight gradient matrix:

∂L∂W(2)=δ(2)(a(1))T=0.034690×[4.00.0]=[0.1387600.000000]\frac{\partial L}{\partial W^{(2)}} = \delta^{(2)} (a^{(1)})^T = 0.034690 \times \begin{bmatrix} 4.0 & 0.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.138760} & \mathbf{0.000000} \end{bmatrix}

Compute the bias gradient:

∂L∂b(2)=δ(2)=+0.034690\frac{\partial L}{\partial b^{(2)}} = \delta^{(2)} = \mathbf{+0.034690}

Attribution Insight: The first output weight W11(2)W^{(2)}_{11} receives a positive gradient of +0.1388+0.1388, indicating that reducing this weight will reduce prediction error. The second weight W12(2)W^{(2)}_{12} receives a gradient of 0.00000.0000 because its incoming hidden activation was zero (a2(1)=0.0a_2^{(1)} = 0.0).


Previous
The Composite Function Chain Rule