Decompose output layer gradients into loss sensitivity, activation slope, and input terms to compute exact credit attribution for output weights.
Now let us apply the chain rule to the output layer of our neural network.
Consider a 2-layer network. The output layer receives a vector of hidden activations a(1)∈Rd1×1, combines them with output weight matrix W(2)∈R1×d1 and output bias b(2)∈R1×1, passes the result through an activation function σ, and evaluates prediction error against ground truth y:
To update the output weights, we must compute the parameter gradient matrix ∂W(2)∂L.
The 3-Factor Gradient Decomposition
By applying the chain rule, the gradient of the loss with respect to output weight Wij(2) decomposes into three chained derivative factors:
∂Wij(2)∂L=Factor 1: Loss Sensitivity∂a(2)∂L⋅Factor 2: Activation Slope∂z(2)∂a(2)⋅Factor 3: Affine Input∂Wij(2)∂z(2)
Let us derive each of these three factors individually:
Factor 1: Loss Sensitivity (∂a(2)∂L)
For standard Mean Squared Error loss L=(a(2)−y)2:
∂a(2)∂L=da(2)d[(a(2)−y)2]=2(a(2)−y)
This term measures the raw prediction error direction and magnitude. If a(2)>y, the loss sensitivity is positive, signaling that the output was too high.
Factor 2: Activation Function Slope (σ′(z(2)))
For the Sigmoid squashing function a(2)=σ(z(2))=1+e−z(2)1, we compute its derivative using the power and chain rules:
Since 1+e−ze−z=1+e−z1+e−z−1=1−σ(z), we obtain the Sigmoid derivative identity:
σ′(z)=σ(z)(1−σ(z))=a(2)(1−a(2))
This factor acts as a sensitivity gate. When the neuron is saturated near 0 or 1, the slope σ′(z) approaches zero, dampening the error signal.
Factor 3: Affine Input Sensitivity (∂W(2)∂z(2))
The linear pre-activation is z(2)=∑jW1j(2)aj(1)+b(2). Differentiating with respect to weight W1j(2) holding all other terms constant:
∂W1j(2)∂z(2)=aj(1)
And differentiating with respect to the bias b(2):
∂b(2)∂z(2)=1
The Output Layer Error Delta (δ(2))
To keep calculations organized and avoid redundant multiplications, deep learning packages Factors 1 and 2 into a single intermediate term called the Error Delta (denoted δ):
In vector notation (using the Hadamard element-wise product ⊙ for multi-output layers):
δ(2)=∂a(2)∂L⊙σ′(z(2))
The error delta δ(2) represents the exact rate of change of the loss with respect to the pre-activation z(2). It quantifies how much total blame arrives at the output layer's linear sum.
Output Parameter Gradients
Once δ(2) is computed, obtaining the parameter gradients requires only multiplying by the incoming inputs:
∂W(2)∂L=δ(2)(a(1))T∂b(2)∂L=δ(2)
Dimensional Consistency Audit Table
To ensure all operations are valid, let us check matrix shapes for our 4→2→1 network:
Variable
Mathematical Meaning
Local Shape (4→2→1)
General Shape (d0→d1→d2)
a(1)
Hidden Layer Activation Vector
R2×1
Rd1×1
(a(1))T
Transposed Hidden Activation
R1×2
R1×d1
z(2)
Output Pre-activation
R1×1
Rd2×1
a(2)
Output Activation Prediction
R1×1
Rd2×1
∂a(2)∂L
Loss Sensitivity Vector
R1×1
Rd2×1
σ′(z(2))
Activation Slope Vector
R1×1
Rd2×1
δ(2)
Output Error Delta
R1×1
Rd2×1
∂W(2)∂L
Output Weight Gradient Matrix
R1×1×R1×2=R1×2
Rd2×1×R1×d1=Rd2×d1
∂b(2)∂L
Output Bias Gradient Vector
R1×1
Rd2×1
The gradient matrix ∂W(2)∂L∈R1×2 matches the exact dimensions of weight matrix W(2)∈R1×2.
Step-by-Step Numerical Walkthrough: 'Die Hard in Space'
Let us compute the exact output layer gradients for our running movie prediction network on 'Die Hard in Space'.
Attribution Insight: The first output weight W11(2) receives a positive gradient of +0.1388, indicating that reducing this weight will reduce prediction error. The second weight W12(2) receives a gradient of 0.0000 because its incoming hidden activation was zero (a2(1)=0.0).