Assembling the Multivariable Gradient Vector hero
Lesson 3Partial Derivatives and Gradients

Assembling the Multivariable Gradient Vector

Assemble individual partial derivatives into a unified gradient vector matching the exact dimensionality of the neural network parameter space.

In Lesson 2, we derived individual scalar partial derivatives (∂L∂w1,∂L∂w2,…,∂L∂wn\frac{\partial L}{\partial w_1}, \frac{\partial L}{\partial w_2}, \dots, \frac{\partial L}{\partial w_n}).

To coordinate parameter adjustments across the entire network simultaneously, we group all of these scalar derivatives into a single unified mathematical object: the Gradient Vector.


The Gradient Vector (∇L\nabla L)

The gradient of a multivariable function L(w)L(w) is denoted by the nabla symbol ∇\nabla (an inverted Greek delta, ∇\nabla).

For a model with nn weights, the gradient vector ∇wL\nabla_w L is the column vector of all nn first-order partial derivatives:

∇wL=[∂L∂w1∂L∂w2⋮∂L∂wn]∈Rn×1\nabla_w L = \begin{bmatrix} \frac{\partial L}{\partial w_1} \\ \frac{\partial L}{\partial w_2} \\ \vdots \\ \frac{\partial L}{\partial w_n} \end{bmatrix} \in \mathbb{R}^{n \times 1}

Equivalently, written as a transposed row vector:

∇wL=[∂L∂w1∂L∂w2⋯∂L∂wn]T\nabla_w L = \begin{bmatrix} \frac{\partial L}{\partial w_1} & \frac{\partial L}{\partial w_2} & \cdots & \frac{\partial L}{\partial w_n} \end{bmatrix}^T

When including the bias parameter bb, the complete parameter vector θ=[w1⋯wnb]T∈R(n+1)×1\theta = \begin{bmatrix} w_1 & \cdots & w_n & b \end{bmatrix}^T \in \mathbb{R}^{(n+1) \times 1} has the full gradient vector:

∇θL=[∂L∂w1⋮∂L∂wn∂L∂b]∈R(n+1)×1\nabla_\theta L = \begin{bmatrix} \frac{\partial L}{\partial w_1} \\ \vdots \\ \frac{\partial L}{\partial w_n} \\ \frac{\partial L}{\partial b} \end{bmatrix} \in \mathbb{R}^{(n+1) \times 1}

Dimension Matching: Structural Symmetry

A fundamental property of the gradient vector is dimension matching:

If w∈Rn×1,then ∇wL∈Rn×1\text{If } w \in \mathbb{R}^{n \times 1}, \quad \text{then } \nabla_w L \in \mathbb{R}^{n \times 1}

Weight Vector (Coordinates in Parameter Space):
  w = [ w_1,  w_2,  w_3,  ...,  w_n ]^T  ∈ ℝ^(n × 1)

Gradient Vector (Sensitivities in Parameter Space):
  ∇_w L = [ ∂L/∂w_1,  ∂L/∂w_2,  ∂L/∂w_3,  ...,  ∂L/∂w_n ]^T  ∈ ℝ^(n × 1)

Every weight parameter wiw_i at index position ii has a corresponding partial derivative ∂L∂wi\frac{\partial L}{\partial w_i} at the exact same index position ii in the gradient vector.

This 1-to-1 dimensional alignment is what makes vector parameter updates possible:

wnew=wold−η∇wLw_{\text{new}} = w_{\text{old}} - \eta \nabla_w L

Because both vectors have identical shapes, we can scale the gradient by a step factor η\eta (the learning rate) and subtract it element-by-element from the weight vector.


4D Gradient Vector Trace: 'Die Hard in Space'

Let's compute the complete 4-dimensional gradient vector by hand for our Netflix greenlight predictor on 'Die Hard in Space'.

1. Input Feature Measurements & Current Parameter State:

  • Input Feature Vector (x∈R4x \in \mathbb{R}^4): x=[x1(Action)x2(Romance)x3(Comedy)x4(Sci-Fi)]=[0.900.100.200.80]x = \begin{bmatrix} x_1 (\text{Action}) \\ x_2 (\text{Romance}) \\ x_3 (\text{Comedy}) \\ x_4 (\text{Sci-Fi}) \end{bmatrix} = \begin{bmatrix} 0.90 \\ 0.10 \\ 0.20 \\ 0.80 \end{bmatrix}

  • Current Weight Vector (w∈R4w \in \mathbb{R}^4) and Bias (b∈Rb \in \mathbb{R}): w=[w1w2w3w4]=[0.800.100.200.70],b=0.05w = \begin{bmatrix} w_1 \\ w_2 \\ w_3 \\ w_4 \end{bmatrix} = \begin{bmatrix} 0.80 \\ 0.10 \\ 0.20 \\ 0.70 \end{bmatrix}, \qquad b = 0.05

2. Forward Linear Prediction & Error:

Evaluate the forward pre-activation prediction:

y^=wTx+b=(0.80)(0.90)+(0.10)(0.10)+(0.20)(0.20)+(0.70)(0.80)+0.05\hat{y} = w^T x + b = (0.80)(0.90) + (0.10)(0.10) + (0.20)(0.20) + (0.70)(0.80) + 0.05 y^=0.72+0.01+0.04+0.56+0.05=1.38\hat{y} = 0.72 + 0.01 + 0.04 + 0.56 + 0.05 = \mathbf{1.38}
  • Target Outcome: The movie flopped in theaters (y=0.00y = 0.00).
  • Scaled Loss Penalty (Lscaled=12(y−y^)2L_{\text{scaled}} = \frac{1}{2}(y - \hat{y})^2): L=12(0.00−1.38)2=12(1.9044)=0.9522L = \frac{1}{2}(0.00 - 1.38)^2 = \frac{1}{2}(1.9044) = \mathbf{0.9522}
  • Output Error Discrepancy: y^−y=1.38−0.00=+1.38\hat{y} - y = 1.38 - 0.00 = \mathbf{+1.38}

3. Calculating Individual Partial Derivatives:

Using our formula ∂L∂wi=(y^−y)xi\frac{\partial L}{\partial w_i} = (\hat{y} - y) x_i and ∂L∂b=(y^−y)\frac{\partial L}{\partial b} = (\hat{y} - y):

  • Action Weight Partial Derivative (w1w_1): ∂L∂w1=(+1.38)⋅x1=(+1.38)(0.90)=+1.242\frac{\partial L}{\partial w_1} = (+1.38) \cdot x_1 = (+1.38)(0.90) = \mathbf{+1.242}

  • Romance Weight Partial Derivative (w2w_2): ∂L∂w2=(+1.38)⋅x2=(+1.38)(0.10)=+0.138\frac{\partial L}{\partial w_2} = (+1.38) \cdot x_2 = (+1.38)(0.10) = \mathbf{+0.138}

  • Comedy Weight Partial Derivative (w3w_3): ∂L∂w3=(+1.38)⋅x3=(+1.38)(0.20)=+0.276\frac{\partial L}{\partial w_3} = (+1.38) \cdot x_3 = (+1.38)(0.20) = \mathbf{+0.276}

  • Sci-Fi Weight Partial Derivative (w4w_4): ∂L∂w4=(+1.38)⋅x4=(+1.38)(0.80)=+1.104\frac{\partial L}{\partial w_4} = (+1.38) \cdot x_4 = (+1.38)(0.80) = \mathbf{+1.104}

  • Bias Partial Derivative (bb): ∂L∂b=(+1.38)⋅(1.00)=+1.380\frac{\partial L}{\partial b} = (+1.38) \cdot (1.00) = \mathbf{+1.380}

4. Assembling the Gradient Vectors:

∇wL=[∂L∂w1∂L∂w2∂L∂w3∂L∂w4]=[+1.242+0.138+0.276+1.104]∈R4×1\nabla_w L = \begin{bmatrix} \frac{\partial L}{\partial w_1} \\ \frac{\partial L}{\partial w_2} \\ \frac{\partial L}{\partial w_3} \\ \frac{\partial L}{\partial w_4} \end{bmatrix} = \begin{bmatrix} +1.242 \\ +0.138 \\ +0.276 \\ +1.104 \end{bmatrix} \in \mathbb{R}^{4 \times 1} ∇θL=[+1.242+0.138+0.276+1.104+1.380]∈R5×1\nabla_\theta L = \begin{bmatrix} +1.242 \\ +0.138 \\ +0.276 \\ +1.104 \\ +1.380 \end{bmatrix} \in \mathbb{R}^{5 \times 1}
ParameterFeature NameInput Feature (xix_i)Current Weight (wiw_i)Partial Derivative (∂L∂wi\frac{\partial L}{\partial w_i})Attribution Ranking
w1w_1Action0.900.900.800.80+1.242+1.242#2 Largest Contributor
w2w_2Romance0.100.100.100.10+0.138+0.138Smallest Contributor
w3w_3Comedy0.200.200.200.20+0.276+0.276Minor Contributor
w4w_4Sci-Fi0.800.800.700.70+1.104+1.104#3 Largest Contributor
bbBias Offset1.001.000.050.05+1.380+1.380#1 Largest Contributor

Attribution Insight: Action (w1w_1) and Sci-Fi (w4w_4) have large positive partial derivatives (+1.242+1.242 and +1.104+1.104) because their input features were heavily present in the movie (0.900.90 and 0.800.80). Romance (w2w_2) has a small derivative (+0.138+0.138) because the screenplay contained almost no romance (0.100.10).

The gradient vector automatically scales credit attribution proportional to each feature's actual contribution to the mistaken prediction.


Dual-Track Bridge: Assembling Vectors

  • Track 1 (Underlying Mechanism): On a 3D mountain landscape, if the North-South slope is ∂z∂x=+3.0\frac{\partial z}{\partial x} = +3.0 and the East-West slope is ∂z∂y=+4.0\frac{\partial z}{\partial y} = +4.0, assembling the 2D gradient vector ∇z=[3.04.0]\nabla z = \begin{bmatrix} 3.0 \\ 4.0 \end{bmatrix} gives a compass heading pointing Northeast.
  • Track 2 (Applied Concept): In the 4D parameter space of our greenlight model, the gradient vector ∇wL=[1.2420.1380.2761.104]T\nabla_w L = \begin{bmatrix} 1.242 & 0.138 & 0.276 & 1.104 \end{bmatrix}^T bundles all four individual genre sensitivities into a single directional pointer.

Previous
Calculating Partial Derivatives of Weights