Compute partial derivatives for individual parameter weights under Mean Squared Error loss by treating all competing weights as static constants.
Now that we understand the freezing principle, we formalize the mathematics of partial differentiation.
The Partial Derivative Symbol: ∂ vs d
In single-variable calculus, we use the straight Latin letter d:
dxdy(Total derivative: x is the only independent variable)
In multivariable calculus, we use the stylized curly symbol ∂ (pronounced "del" or "partial"):
∂wi∂L(Partial derivative: differentiate with respect to wi, holding all other inputs constant)
The symbol ∂ serves as an explicit mathematical flag: "This system has multiple independent inputs. We are measuring sensitivity to one specific variable while holding every other variable strictly frozen."
Formal Limit Definition of the Partial Derivative
For a multivariable function f(w1,w2,…,wn), the formal partial derivative with respect to parameter wi is defined as:
Notice that in the numerator, only the i-th variable receives the perturbation Δwi. Every other parameter (w1,…,wi−1,wi+1,…,wn) remains completely identical in both terms.
Step-by-Step Differentiation: The Affine Linear Sum
Let's compute the partial derivatives of the fundamental neural network operation: the affine linear pre-activation sum z:
z=w1x1+w2x2+⋯+wnxn+b=j=1∑nwjxj+b
Let's differentiate z with respect to weight w1, holding all other weights (w2,…,wn) and bias b constant:
First term (w1x1): With respect to w1, the input feature x1 is a constant coefficient multiplier. Applying the linear multiplier rule (dxd[cx]=c):
∂w1∂[w1x1]=x1
Competing weight terms (w2x2,…,wnxn): None of these terms contain w1. Because w2,…,wn and x2,…,xn are treated as constants, their product is a constant number. The derivative of any constant is zero:
∂w1∂[w2x2]=0,…,∂w1∂[wnxn]=0
Bias term (b): The bias is a constant additive offset with no w1 term:
∂w1∂[b]=0
Sum the results:∂w1∂z=x1+0+⋯+0+0=x1
By identical logic, the partial derivative of the linear sum with respect to any arbitrary weight wi is simply its corresponding feature measurement:
∂wi∂z=xi
Differentiating with Respect to Bias b:
Now let's differentiate z with respect to the bias parameter b, holding all weights (w1,…,wn) constant:
Every weight product wjxj contains no b term, so their derivatives are all zero:
∂b∂[wjxj]=0for all j∈{1,…,n}
The bias term b has a linear coefficient of 1:
∂b∂[b]=1
Summing the terms:
∂b∂z=0+0+⋯+0+1=1
Summary of Linear Sum Sensitivities:
∂z / ∂w_i = x_i (The sensitivity of pre-activation to a weight is the feature measurement itself)
∂z / ∂b = 1 (The sensitivity of pre-activation to bias is always exactly 1)
Step-by-Step Partial Differentiation of Linear MSE Loss
Now let's connect parameter weights directly to the final Mean Squared Error loss.
Consider a single-layer linear model where predicted output equals pre-activation (y^=z=w1x1+w2x2+b), evaluated against ground-truth target y:
L=(y−y^)2=(y−(w1x1+w2x2+b))2
Let's compute the partial derivative ∂w1∂L using algebraic expansion and term-by-term differentiation:
Method 1: Expanding the Quadratic Expression
Expand the squared loss formula:
L=y2−2y(w1x1+w2x2+b)+(w1x1+w2x2+b)2
Now differentiate each part with respect to w1, treating y,w2,b,x1,x2 as constants:
Target constant term (y2):∂w1∂[y2]=0
Cross-product term (−2y(w1x1+w2x2+b)):∂w1∂[−2yw1x1−2yw2x2−2yb]=−2yx1−0−0=−2yx1
Squared sum term ((w1x1+w2x2+b)2=y^2):
Applying the power rule and inner linear derivative:
∂w1∂[(w1x1+w2x2+b)2]=2(w1x1+w2x2+b)⋅∂w1∂[w1x1+w2x2+b]=2y^⋅x1
Combine the derivatives:∂w1∂L=−2yx1+2y^x1=2(y^−y)x1
Under the 21 Scaled Loss Convention:
For the standard calculus-scaled Mean Squared Error Lscaled=21(y−y^)2:
∂w1∂Lscaled=21⋅2(y^−y)x1=(y^−y)x1
The Fundamental Error Attribution Factoring:
Notice how the partial derivative naturally decomposes into two clean components:
∂wi∂L=Output Loss Derivative ∂y^∂L2(y^−y)⋅Input Feature Derivative ∂wi∂y^xi
The first term, 2(y^−y), is the output prediction error (how wrong the model's prediction was).
The second term, xi, is the feature measurement (how active that specific input was during the forward pass).
If a feature was absent (xi=0), its weight partial derivative is ∂wi∂L=0. A weight cannot be blamed for a prediction error if its input feature was never active.