Calculating Partial Derivatives of Weights hero
Lesson 2Partial Derivatives and Gradients

Calculating Partial Derivatives of Weights

Compute partial derivatives for individual parameter weights under Mean Squared Error loss by treating all competing weights as static constants.

Now that we understand the freezing principle, we formalize the mathematics of partial differentiation.


The Partial Derivative Symbol: ∂\partial vs dd

In single-variable calculus, we use the straight Latin letter dd:

dydx(Total derivative: x is the only independent variable)\frac{dy}{dx} \qquad \text{(Total derivative: $x$ is the only independent variable)}

In multivariable calculus, we use the stylized curly symbol ∂\partial (pronounced "del" or "partial"):

∂L∂wi(Partial derivative: differentiate with respect to wi, holding all other inputs constant)\frac{\partial L}{\partial w_i} \qquad \text{(Partial derivative: differentiate with respect to $w_i$, holding all other inputs constant)}

The symbol ∂\partial serves as an explicit mathematical flag: "This system has multiple independent inputs. We are measuring sensitivity to one specific variable while holding every other variable strictly frozen."


Formal Limit Definition of the Partial Derivative

For a multivariable function f(w1,w2,…,wn)f(w_1, w_2, \dots, w_n), the formal partial derivative with respect to parameter wiw_i is defined as:

∂f∂wi=lim⁡Δwi→0f(w1,…,wi+Δwi,…,wn)−f(w1,…,wi,…,wn)Δwi\frac{\partial f}{\partial w_i} = \lim_{\Delta w_i \to 0} \frac{f(w_1, \dots, w_i + \Delta w_i, \dots, w_n) - f(w_1, \dots, w_i, \dots, w_n)}{\Delta w_i}

Notice that in the numerator, only the ii-th variable receives the perturbation Δwi\Delta w_i. Every other parameter (w1,…,wi−1,wi+1,…,wnw_1, \dots, w_{i-1}, w_{i+1}, \dots, w_n) remains completely identical in both terms.


Step-by-Step Differentiation: The Affine Linear Sum

Let's compute the partial derivatives of the fundamental neural network operation: the affine linear pre-activation sum zz:

z=w1x1+w2x2+⋯+wnxn+b=∑j=1nwjxj+bz = w_1 x_1 + w_2 x_2 + \dots + w_n x_n + b = \sum_{j=1}^n w_j x_j + b

Let's differentiate zz with respect to weight w1w_1, holding all other weights (w2,…,wnw_2, \dots, w_n) and bias bb constant:

  1. First term (w1x1w_1 x_1): With respect to w1w_1, the input feature x1x_1 is a constant coefficient multiplier. Applying the linear multiplier rule (ddx[cx]=c\frac{d}{dx}[cx] = c): ∂∂w1[w1x1]=x1\frac{\partial}{\partial w_1}[w_1 x_1] = x_1

  2. Competing weight terms (w2x2,…,wnxnw_2 x_2, \dots, w_n x_n): None of these terms contain w1w_1. Because w2,…,wnw_2, \dots, w_n and x2,…,xnx_2, \dots, x_n are treated as constants, their product is a constant number. The derivative of any constant is zero: ∂∂w1[w2x2]=0,…,∂∂w1[wnxn]=0\frac{\partial}{\partial w_1}[w_2 x_2] = 0, \quad \dots, \quad \frac{\partial}{\partial w_1}[w_n x_n] = 0

  3. Bias term (bb): The bias is a constant additive offset with no w1w_1 term: ∂∂w1[b]=0\frac{\partial}{\partial w_1}[b] = 0

  4. Sum the results: ∂z∂w1=x1+0+⋯+0+0=x1\frac{\partial z}{\partial w_1} = x_1 + 0 + \dots + 0 + 0 = \mathbf{x_1}

By identical logic, the partial derivative of the linear sum with respect to any arbitrary weight wiw_i is simply its corresponding feature measurement:

∂z∂wi=xi\frac{\partial z}{\partial w_i} = x_i

Differentiating with Respect to Bias bb:

Now let's differentiate zz with respect to the bias parameter bb, holding all weights (w1,…,wnw_1, \dots, w_n) constant:

  1. Every weight product wjxjw_j x_j contains no bb term, so their derivatives are all zero: ∂∂b[wjxj]=0for all j∈{1,…,n}\frac{\partial}{\partial b}[w_j x_j] = 0 \quad \text{for all } j \in \{1, \dots, n\}

  2. The bias term bb has a linear coefficient of 11: ∂∂b[b]=1\frac{\partial}{\partial b}[b] = 1

  3. Summing the terms: ∂z∂b=0+0+⋯+0+1=1\frac{\partial z}{\partial b} = 0 + 0 + \dots + 0 + 1 = \mathbf{1}

Summary of Linear Sum Sensitivities:
  ∂z / ∂w_i = x_i   (The sensitivity of pre-activation to a weight is the feature measurement itself)
  ∂z / ∂b   = 1     (The sensitivity of pre-activation to bias is always exactly 1)

Step-by-Step Partial Differentiation of Linear MSE Loss

Now let's connect parameter weights directly to the final Mean Squared Error loss.

Consider a single-layer linear model where predicted output equals pre-activation (y^=z=w1x1+w2x2+b\hat{y} = z = w_1 x_1 + w_2 x_2 + b), evaluated against ground-truth target yy:

L=(y−y^)2=(y−(w1x1+w2x2+b))2L = (y - \hat{y})^2 = \left(y - (w_1 x_1 + w_2 x_2 + b)\right)^2

Let's compute the partial derivative ∂L∂w1\frac{\partial L}{\partial w_1} using algebraic expansion and term-by-term differentiation:

Method 1: Expanding the Quadratic Expression

Expand the squared loss formula:

L=y2−2y(w1x1+w2x2+b)+(w1x1+w2x2+b)2L = y^2 - 2y(w_1 x_1 + w_2 x_2 + b) + (w_1 x_1 + w_2 x_2 + b)^2

Now differentiate each part with respect to w1w_1, treating y,w2,b,x1,x2y, w_2, b, x_1, x_2 as constants:

  1. Target constant term (y2y^2): ∂∂w1[y2]=0\frac{\partial}{\partial w_1}[y^2] = 0

  2. Cross-product term (−2y(w1x1+w2x2+b)-2y(w_1 x_1 + w_2 x_2 + b)): ∂∂w1[−2yw1x1−2yw2x2−2yb]=−2yx1−0−0=−2yx1\frac{\partial}{\partial w_1}[-2y w_1 x_1 - 2y w_2 x_2 - 2yb] = -2y x_1 - 0 - 0 = -2y x_1

  3. Squared sum term ((w1x1+w2x2+b)2=y^2(w_1 x_1 + w_2 x_2 + b)^2 = \hat{y}^2): Applying the power rule and inner linear derivative: ∂∂w1[(w1x1+w2x2+b)2]=2(w1x1+w2x2+b)⋅∂∂w1[w1x1+w2x2+b]=2y^⋅x1\frac{\partial}{\partial w_1}[(w_1 x_1 + w_2 x_2 + b)^2] = 2(w_1 x_1 + w_2 x_2 + b) \cdot \frac{\partial}{\partial w_1}[w_1 x_1 + w_2 x_2 + b] = 2\hat{y} \cdot x_1

  4. Combine the derivatives: ∂L∂w1=−2yx1+2y^x1=2(y^−y)x1\frac{\partial L}{\partial w_1} = -2y x_1 + 2\hat{y} x_1 = \mathbf{2(\hat{y} - y) x_1}

Under the 12\frac{1}{2} Scaled Loss Convention:

For the standard calculus-scaled Mean Squared Error Lscaled=12(y−y^)2L_{\text{scaled}} = \frac{1}{2}(y - \hat{y})^2:

∂Lscaled∂w1=12⋅2(y^−y)x1=(y^−y)x1\frac{\partial L_{\text{scaled}}}{\partial w_1} = \frac{1}{2} \cdot 2(\hat{y} - y) x_1 = \mathbf{(\hat{y} - y) x_1}

The Fundamental Error Attribution Factoring:

Notice how the partial derivative naturally decomposes into two clean components:

∂L∂wi=2(y^−y)⏟Output Loss Derivative ∂L∂y^⋅xi⏟Input Feature Derivative ∂y^∂wi\frac{\partial L}{\partial w_i} = \underbrace{2(\hat{y} - y)}_{\text{Output Loss Derivative } \frac{\partial L}{\partial \hat{y}}} \cdot \underbrace{x_i}_{\text{Input Feature Derivative } \frac{\partial \hat{y}}{\partial w_i}}
  • The first term, 2(y^−y)2(\hat{y} - y), is the output prediction error (how wrong the model's prediction was).
  • The second term, xix_i, is the feature measurement (how active that specific input was during the forward pass).

If a feature was absent (xi=0x_i = 0), its weight partial derivative is ∂L∂wi=0\frac{\partial L}{\partial w_i} = 0. A weight cannot be blamed for a prediction error if its input feature was never active.


Previous
Multivariable Systems and Freezing