Assemble individual partial derivatives into a unified gradient vector matching the exact dimensionality of the neural network parameter space.
In Lesson 2, we derived individual scalar partial derivatives (∂w1∂L,∂w2∂L,…,∂wn∂L).
To coordinate parameter adjustments across the entire network simultaneously, we group all of these scalar derivatives into a single unified mathematical object: the Gradient Vector.
The Gradient Vector (∇L)
The gradient of a multivariable function L(w) is denoted by the nabla symbol ∇ (an inverted Greek delta, ∇).
For a model with n weights, the gradient vector ∇wL is the column vector of all n first-order partial derivatives:
∇wL=∂w1∂L∂w2∂L⋮∂wn∂L∈Rn×1
Equivalently, written as a transposed row vector:
∇wL=[∂w1∂L∂w2∂L⋯∂wn∂L]T
When including the bias parameter b, the complete parameter vector θ=[w1⋯wnb]T∈R(n+1)×1 has the full gradient vector:
∇θL=∂w1∂L⋮∂wn∂L∂b∂L∈R(n+1)×1
Dimension Matching: Structural Symmetry
A fundamental property of the gradient vector is dimension matching:
If w∈Rn×1,then ∇wL∈Rn×1
Weight Vector (Coordinates in Parameter Space):
w = [ w_1, w_2, w_3, ..., w_n ]^T ∈ ℝ^(n × 1)
Gradient Vector (Sensitivities in Parameter Space):
∇_w L = [ ∂L/∂w_1, ∂L/∂w_2, ∂L/∂w_3, ..., ∂L/∂w_n ]^T ∈ ℝ^(n × 1)
Every weight parameter wi at index position i has a corresponding partial derivative ∂wi∂L at the exact same index position i in the gradient vector.
This 1-to-1 dimensional alignment is what makes vector parameter updates possible:
wnew=wold−η∇wL
Because both vectors have identical shapes, we can scale the gradient by a step factor η (the learning rate) and subtract it element-by-element from the weight vector.
4D Gradient Vector Trace: 'Die Hard in Space'
Let's compute the complete 4-dimensional gradient vector by hand for our Netflix greenlight predictor on 'Die Hard in Space'.
1. Input Feature Measurements & Current Parameter State:
Attribution Insight:
Action (w1) and Sci-Fi (w4) have large positive partial derivatives (+1.242 and +1.104) because their input features were heavily present in the movie (0.90 and 0.80). Romance (w2) has a small derivative (+0.138) because the screenplay contained almost no romance (0.10).
The gradient vector automatically scales credit attribution proportional to each feature's actual contribution to the mistaken prediction.
Dual-Track Bridge: Assembling Vectors
Track 1 (Underlying Mechanism): On a 3D mountain landscape, if the North-South slope is ∂x∂z=+3.0 and the East-West slope is ∂y∂z=+4.0, assembling the 2D gradient vector ∇z=[3.04.0] gives a compass heading pointing Northeast.
Track 2 (Applied Concept): In the 4D parameter space of our greenlight model, the gradient vector ∇wL=[1.2420.1380.2761.104]T bundles all four individual genre sensitivities into a single directional pointer.