Hidden Layer Error Backpropagation hero
Lesson 3The Chain Rule and Backpropagation

Hidden Layer Error Backpropagation

Propagate error signals backward through hidden layers by multiplying downstream deltas by transposed weights and intermediate activation slopes.

We now have the output layer error delta δ(2)\delta^{(2)}. But how do we determine how much each hidden weight Wij(1)W^{(1)}_{ij} contributed to the error?


The Hidden Layer Credit Assignment Problem

Hidden neurons have no external ground-truth labels. The world does not tell the network what the Summer Popcorn Flick neuron (h1h_1) or the Rom-Com neuron (h2h_2) should have produced.

To solve this credit assignment problem, we propagate the downstream error delta δ(2)\delta^{(2)} backward through the network's weights.

[ Forward Pass ]:  Inputs x ──► Hidden Layer (z^(1), a^(1)) ──► Output Layer (z^(2), a^(2)) ──► Loss L
                                                                                                  │
[ Backward Pass ]: ∂L/∂W^(1) ◄── Hidden Delta δ^(1) ◄── (W^(2))^T ◄── Output Delta δ^(2) ◄────────┘

Backpropagating Error Through Transposed Weights

To evaluate how sensitive the final loss is to a hidden activation aj(1)a_j^{(1)}, apply the multivariable chain rule across all output pre-activations zk(2)z_k^{(2)} that depend on aj(1)a_j^{(1)}:

∂L∂aj(1)=∑k=1d2∂L∂zk(2)∂zk(2)∂aj(1)=∑k=1d2δk(2)Wkj(2)\frac{\partial L}{\partial a_j^{(1)}} = \sum_{k=1}^{d_2} \frac{\partial L}{\partial z_k^{(2)}} \frac{\partial z_k^{(2)}}{\partial a_j^{(1)}} = \sum_{k=1}^{d_2} \delta_k^{(2)} W^{(2)}_{kj}

In matrix-vector notation, this sum is computed by multiplying the downstream delta vector by the transposed weight matrix:

∂L∂a(1)=(W(2))Tδ(2)\frac{\partial L}{\partial a^{(1)}} = (W^{(2)})^T \delta^{(2)}

Why Do We Transpose the Weight Matrix?

  • In the forward pass, W(2)∈Rd2×d1W^{(2)} \in \mathbb{R}^{d_2 \times d_1} maps hidden activations a(1)∈Rd1×1a^{(1)} \in \mathbb{R}^{d_1 \times 1} forward to output pre-activations z(2)∈Rd2×1z^{(2)} \in \mathbb{R}^{d_2 \times 1}.
  • In the backward pass, we must map output error deltas δ(2)∈Rd2×1\delta^{(2)} \in \mathbb{R}^{d_2 \times 1} backward to the hidden layer dimension Rd1×1\mathbb{R}^{d_1 \times 1}.
  • Transposing flips the matrix shape to (W(2))T∈Rd1×d2(W^{(2)})^T \in \mathbb{R}^{d_1 \times d_2}, making the matrix-vector multiplication (W(2))Tδ(2)(W^{(2)})^T \delta^{(2)} dimensionally consistent: Rd1×d2×Rd2×1=Rd1×1\mathbb{R}^{d_1 \times d_2} \times \mathbb{R}^{d_2 \times 1} = \mathbb{R}^{d_1 \times 1}

The Hidden Layer Error Delta (δ(1)\delta^{(1)})

To convert the activation sensitivity ∂L∂a(1)\frac{\partial L}{\partial a^{(1)}} into the pre-activation error delta δ(1)≡∂L∂z(1)\delta^{(1)} \equiv \frac{\partial L}{\partial z^{(1)}}, we apply the chain rule across the hidden activation function ff:

δj(1)=∂L∂aj(1)⋅∂aj(1)∂zj(1)=(∑kδk(2)Wkj(2))⋅f′(zj(1))\delta_j^{(1)} = \frac{\partial L}{\partial a_j^{(1)}} \cdot \frac{\partial a_j^{(1)}}{\partial z_j^{(1)}} = \left( \sum_k \delta_k^{(2)} W^{(2)}_{kj} \right) \cdot f'(z_j^{(1)})

In compact vector notation:

δ(1)=((W(2))Tδ(2))⊙f′(z(1))\mathbf{\delta^{(1)} = \left( (W^{(2)})^T \delta^{(2)} \right) \odot f'(z^{(1)})}

Where:

  • For ReLU activation (f(z)=max⁡(0,z)f(z) = \max(0, z)): f′(z)={1.0if z>00.0if z≤0f'(z) = \begin{cases} 1.0 & \text{if } z > 0 \\ 0.0 & \text{if } z \le 0 \end{cases}
  • For Sigmoid activation (f(z)=σ(z)f(z) = \sigma(z)): f′(z)=σ(z)(1−σ(z))=a(1)⊙(1−a(1))f'(z) = \sigma(z)(1 - \sigma(z)) = a^{(1)} \odot (1 - a^{(1)})

Hidden Layer Parameter Gradients

Once the hidden delta vector δ(1)∈Rd1×1\delta^{(1)} \in \mathbb{R}^{d_1 \times 1} is established, we compute the gradients for the first-layer weights W(1)∈Rd1×d0W^{(1)} \in \mathbb{R}^{d_1 \times d_0} and biases b(1)∈Rd1×1b^{(1)} \in \mathbb{R}^{d_1 \times 1} using the outer product with the input feature vector x∈Rd0×1x \in \mathbb{R}^{d_0 \times 1}:

∂L∂W(1)=δ(1)xT\mathbf{\frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T} ∂L∂b(1)=δ(1)\mathbf{\frac{\partial L}{\partial b^{(1)}} = \delta^{(1)}}

Dimensional Consistency Audit for Layer 1

Mathematical ObjectVariableDimension (4→2→14 \to 2 \to 1)Dimension (d0→d1→d2d_0 \to d_1 \to d_2)
Input Feature VectorxxR4×1\mathbb{R}^{4 \times 1}Rd0×1\mathbb{R}^{d_0 \times 1}
Transposed Input VectorxTx^TR1×4\mathbb{R}^{1 \times 4}R1×d0\mathbb{R}^{1 \times d_0}
Transposed Output Weight Matrix(W(2))T(W^{(2)})^TR2×1\mathbb{R}^{2 \times 1}Rd1×d2\mathbb{R}^{d_1 \times d_2}
Backpropagated Error Signal(W(2))Tδ(2)(W^{(2)})^T \delta^{(2)}R2×1\mathbb{R}^{2 \times 1}Rd1×1\mathbb{R}^{d_1 \times 1}
Hidden Activation Slope Vectorf′(z(1))f'(z^{(1)})R2×1\mathbb{R}^{2 \times 1}Rd1×1\mathbb{R}^{d_1 \times 1}
Hidden Error Delta Vectorδ(1)\delta^{(1)}R2×1\mathbb{R}^{2 \times 1}Rd1×1\mathbb{R}^{d_1 \times 1}
Hidden Weight Gradient Matrix∂L∂W(1)=δ(1)xT\frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^TR2×1×R1×4=R2×4\mathbb{R}^{2 \times 1} \times \mathbb{R}^{1 \times 4} = \mathbf{\mathbb{R}^{2 \times 4}}Rd1×1×R1×d0=Rd1×d0\mathbb{R}^{d_1 \times 1} \times \mathbb{R}^{1 \times d_0} = \mathbf{\mathbb{R}^{d_1 \times d_0}}
Hidden Bias Gradient Vector∂L∂b(1)=δ(1)\frac{\partial L}{\partial b^{(1)}} = \delta^{(1)}R2×1\mathbf{\mathbb{R}^{2 \times 1}}Rd1×1\mathbf{\mathbb{R}^{d_1 \times 1}}

Step-by-Step Numerical Walkthrough: 'Die Hard in Space'

Let us continue our backward pass on 'Die Hard in Space' to compute the hidden layer gradients.

  • Input feature vector: x=[1.00.00.51.0](Action Measurement)(Romance Measurement)(Comedy Measurement)(Sci-Fi Measurement)x = \begin{bmatrix} 1.0 \\ 0.0 \\ 0.5 \\ 1.0 \end{bmatrix} \begin{matrix} \text{(Action Measurement)} \\ \text{(Romance Measurement)} \\ \text{(Comedy Measurement)} \\ \text{(Sci-Fi Measurement)} \end{matrix}
  • Layer 1 pre-activations: z(1)=[+4.0−3.5]z^{(1)} = \begin{bmatrix} +4.0 \\ -3.5 \end{bmatrix}
  • Layer 2 weights: W(2)=[1.51.0]W^{(2)} = \begin{bmatrix} 1.5 & 1.0 \end{bmatrix}
  • Output error delta: δ(2)=0.034690\delta^{(2)} = 0.034690

Step 1: Project Error Through Transposed Weights

(W(2))Tδ(2)=[1.51.0](0.034690)=[1.5×0.0346901.0×0.034690]=[0.0520350.034690](W^{(2)})^T \delta^{(2)} = \begin{bmatrix} 1.5 \\ 1.0 \end{bmatrix} (0.034690) = \begin{bmatrix} 1.5 \times 0.034690 \\ 1.0 \times 0.034690 \end{bmatrix} = \begin{bmatrix} \mathbf{0.052035} \\ \mathbf{0.034690} \end{bmatrix}

Step 2: Compute Hidden Activation Slopes

For the ReLU activation function:

f′(z(1))=[ReLU′(+4.0)ReLU′(−3.5)]=[1.00.0]f'(z^{(1)}) = \begin{bmatrix} \text{ReLU}'(+4.0) \\ \text{ReLU}'(-3.5) \end{bmatrix} = \begin{bmatrix} \mathbf{1.0} \\ \mathbf{0.0} \end{bmatrix}

Step 3: Compute Hidden Error Delta δ(1)\delta^{(1)}

Multiply element-wise by the activation slopes:

δ(1)=[0.0520350.034690]⊙[1.00.0]=[0.0520350.000000]\delta^{(1)} = \begin{bmatrix} 0.052035 \\ 0.034690 \end{bmatrix} \odot \begin{bmatrix} 1.0 \\ 0.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.052035} \\ \mathbf{0.000000} \end{bmatrix}

Pedagogical Insight: Because Hidden Neuron 2 (Rom-Com) had a negative pre-activation (z2(1)=−3.5z_2^{(1)} = -3.5), its ReLU derivative is 0.00.0. This completely zeroes out its error delta (δ2(1)=0.0\delta_2^{(1)} = 0.0). The network attributes 100% of the hidden blame to Neuron 1 (Summer Popcorn Flick), protecting dormant neurons from unwanted parameter corruption.

Step 4: Compute Hidden Layer Parameter Gradients

Compute the outer product with xT=[1.00.00.51.0]x^T = \begin{bmatrix} 1.0 & 0.0 & 0.5 & 1.0 \end{bmatrix}:

∂L∂W(1)=δ(1)xT=[0.0520350.000000][1.00.00.51.0]\frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T = \begin{bmatrix} 0.052035 \\ 0.000000 \end{bmatrix} \begin{bmatrix} 1.0 & 0.0 & 0.5 & 1.0 \end{bmatrix} ∂L∂W(1)=[(0.052035)(1.0)(0.052035)(0.0)(0.052035)(0.5)(0.052035)(1.0)(0)(1.0)(0)(0.0)(0)(0.5)(0)(1.0)]\frac{\partial L}{\partial W^{(1)}} = \begin{bmatrix} (0.052035)(1.0) & (0.052035)(0.0) & (0.052035)(0.5) & (0.052035)(1.0) \\ (0)(1.0) & (0)(0.0) & (0)(0.5) & (0)(1.0) \end{bmatrix} ∂L∂W(1)=[0.0520350.0000000.0260180.0520350.0000000.0000000.0000000.000000]\frac{\partial L}{\partial W^{(1)}} = \begin{bmatrix} \mathbf{0.052035} & \mathbf{0.000000} & \mathbf{0.026018} & \mathbf{0.052035} \\ \mathbf{0.000000} & \mathbf{0.000000} & \mathbf{0.000000} & \mathbf{0.000000} \end{bmatrix}

Compute the bias gradient:

∂L∂b(1)=δ(1)=[0.0520350.000000]\frac{\partial L}{\partial b^{(1)}} = \delta^{(1)} = \begin{bmatrix} \mathbf{0.052035} \\ \mathbf{0.000000} \end{bmatrix}

Attribution Complete: Backpropagation has traced the studio flop error all the way back to the screenplay's raw features. The Action weight (W11(1)W^{(1)}_{11}) and Sci-Fi weight (W14(1)W^{(1)}_{14}) receive the largest corrective gradients (+0.0520+0.0520), while Romance receives zero gradient because it was absent from the script (x2=0.0x_2 = 0.0).


Previous
Output Layer Error Attribution Math