Propagate error signals backward through hidden layers by multiplying downstream deltas by transposed weights and intermediate activation slopes.
We now have the output layer error delta δ(2). But how do we determine how much each hidden weight Wij(1) contributed to the error?
The Hidden Layer Credit Assignment Problem
Hidden neurons have no external ground-truth labels. The world does not tell the network what the Summer Popcorn Flick neuron (h1) or the Rom-Com neuron (h2) should have produced.
To solve this credit assignment problem, we propagate the downstream error delta δ(2) backward through the network's weights.
To evaluate how sensitive the final loss is to a hidden activation aj(1), apply the multivariable chain rule across all output pre-activations zk(2) that depend on aj(1):
In matrix-vector notation, this sum is computed by multiplying the downstream delta vector by the transposed weight matrix:
∂a(1)∂L=(W(2))Tδ(2)
Why Do We Transpose the Weight Matrix?
In the forward pass, W(2)∈Rd2×d1 maps hidden activations a(1)∈Rd1×1 forward to output pre-activations z(2)∈Rd2×1.
In the backward pass, we must map output error deltas δ(2)∈Rd2×1 backward to the hidden layer dimension Rd1×1.
Transposing flips the matrix shape to (W(2))T∈Rd1×d2, making the matrix-vector multiplication (W(2))Tδ(2) dimensionally consistent:
Rd1×d2×Rd2×1=Rd1×1
The Hidden Layer Error Delta (δ(1))
To convert the activation sensitivity ∂a(1)∂L into the pre-activation error delta δ(1)≡∂z(1)∂L, we apply the chain rule across the hidden activation function f:
For ReLU activation (f(z)=max(0,z)):
f′(z)={1.00.0if z>0if z≤0
For Sigmoid activation (f(z)=σ(z)):
f′(z)=σ(z)(1−σ(z))=a(1)⊙(1−a(1))
Hidden Layer Parameter Gradients
Once the hidden delta vector δ(1)∈Rd1×1 is established, we compute the gradients for the first-layer weights W(1)∈Rd1×d0 and biases b(1)∈Rd1×1 using the outer product with the input feature vector x∈Rd0×1:
∂W(1)∂L=δ(1)xT∂b(1)∂L=δ(1)
Dimensional Consistency Audit for Layer 1
Mathematical Object
Variable
Dimension (4→2→1)
Dimension (d0→d1→d2)
Input Feature Vector
x
R4×1
Rd0×1
Transposed Input Vector
xT
R1×4
R1×d0
Transposed Output Weight Matrix
(W(2))T
R2×1
Rd1×d2
Backpropagated Error Signal
(W(2))Tδ(2)
R2×1
Rd1×1
Hidden Activation Slope Vector
f′(z(1))
R2×1
Rd1×1
Hidden Error Delta Vector
δ(1)
R2×1
Rd1×1
Hidden Weight Gradient Matrix
∂W(1)∂L=δ(1)xT
R2×1×R1×4=R2×4
Rd1×1×R1×d0=Rd1×d0
Hidden Bias Gradient Vector
∂b(1)∂L=δ(1)
R2×1
Rd1×1
Step-by-Step Numerical Walkthrough: 'Die Hard in Space'
Let us continue our backward pass on 'Die Hard in Space' to compute the hidden layer gradients.
Pedagogical Insight: Because Hidden Neuron 2 (Rom-Com) had a negative pre-activation (z2(1)=−3.5), its ReLU derivative is 0.0. This completely zeroes out its error delta (δ2(1)=0.0). The network attributes 100% of the hidden blame to Neuron 1 (Summer Popcorn Flick), protecting dormant neurons from unwanted parameter corruption.
Step 4: Compute Hidden Layer Parameter Gradients
Compute the outer product with xT=[1.00.00.51.0]:
Attribution Complete: Backpropagation has traced the studio flop error all the way back to the screenplay's raw features. The Action weight (W11(1)) and Sci-Fi weight (W14(1)) receive the largest corrective gradients (+0.0520), while Romance receives zero gradient because it was absent from the script (x2=0.0).