Partial Derivatives & Gradient Vectors In Practice hero
LaboratoryPartial Derivatives and Gradients

Partial Derivatives & Gradient Vectors In Practice

Master multivariable parameter isolation, partial derivative calculations, gradient vector assembly, and steepest descent trajectories through manual calculations.

Work through each problem on paper before expanding the solution details.


Interactive Sandbox: 2D Loss Surface & Gradient Vector Visualizer

Before proceeding to manual calculations, explore the geometric behavior of 2D loss bowls:

  • Interactive Controls:
    • Draggable parameter cursor (w1,w2)(w_1, w_2) on a 2D contour map and 3D loss surface.
    • Parameter Isolation Toggle: Lock w2w_2 to view a 1D parabolic cross-section along the w1w_1 axis.
    • Live vector component readouts: ∂L∂w1\frac{\partial L}{\partial w_1}, ∂L∂w2\frac{\partial L}{\partial w_2}, and magnitude ∥∇L∥\|\nabla L\|.
    • Rendered vector arrows: Uphill Gradient (+∇L+\nabla L, red) and Downhill Descent (−∇L-\nabla L, green).
  • Curated Presets:
    1. Anisotropic Loss Bowl: L(w1,w2)=4(w1−2)2+(w2−3)2L(w_1, w_2) = 4(w_1 - 2)^2 + (w_2 - 3)^2 (steeper curvature along w1w_1 than w2w_2).
    2. Flop Error Coordinate: High-error operating point at (w1,w2)=(5.0,4.0)(w_1, w_2) = (5.0, 4.0) demonstrating large gradient magnitude.
    3. Optimal Valley Minimum: Global minimum at (w1∗,w2∗)=(2.0,3.0)(w_1^*, w_2^*) = (2.0, 3.0) where ∇L=[00]\nabla L = \begin{bmatrix} 0 \\ 0 \end{bmatrix} and ∥∇L∥=0\|\nabla L\| = 0.
  • Guided Challenge Cue:

    "Challenge: Lock w2w_2 and adjust w1w_1 to find the 1D slice minimum, then release w2w_2 and observe how the true 2D gradient points along a combined descent trajectory perpendicular to the contour lines."


Part 1: The Underlying Mechanics Drill

Problem 1: Partial Derivative with Respect to xx

Given the multivariable polynomial function:

f(x,y)=3x2y+4x−5y2f(x, y) = 3x^2 y + 4x - 5y^2

Compute the partial derivative ∂f∂x\frac{\partial f}{\partial x} by treating yy as a constant.

Reveal Solution

Step 1: Differentiate each term with respect to xx, treating yy as a constant:

  1. Term 1 (3x2y3x^2 y): The factor 3y3y acts as a constant multiplier on x2x^2: ∂∂x[3y⋅x2]=3y⋅(2x)=6xy\frac{\partial}{\partial x}[3y \cdot x^2] = 3y \cdot (2x) = 6xy

  2. Term 2 (4x4x): Linear multiplier rule: ∂∂x[4x]=4\frac{\partial}{\partial x}[4x] = 4

  3. Term 3 (−5y2-5y^2): Contains no xx variable. Since yy is constant, −5y2-5y^2 is constant: ∂∂x[−5y2]=0\frac{\partial}{\partial x}[-5y^2] = 0

Step 2: Combine the terms:

∂f∂x=6xy+4+0=6xy+4\frac{\partial f}{\partial x} = 6xy + 4 + 0 = \mathbf{6xy + 4}

Problem 2: Partial Derivative with Respect to yy

Using the same function:

f(x,y)=3x2y+4x−5y2f(x, y) = 3x^2 y + 4x - 5y^2

Compute the partial derivative ∂f∂y\frac{\partial f}{\partial y} by treating xx as a constant.

Reveal Solution

Step 1: Differentiate each term with respect to yy, treating xx as a constant:

  1. Term 1 (3x2y3x^2 y): The factor 3x23x^2 acts as a constant multiplier on yy: ∂∂y[3x2⋅y]=3x2⋅(1)=3x2\frac{\partial}{\partial y}[3x^2 \cdot y] = 3x^2 \cdot (1) = 3x^2

  2. Term 2 (4x4x): Contains no yy variable. Since xx is constant, 4x4x is constant: ∂∂y[4x]=0\frac{\partial}{\partial y}[4x] = 0

  3. Term 3 (−5y2-5y^2): Standard power rule on yy: ∂∂y[−5y2]=−5⋅(2y)=−10y\frac{\partial}{\partial y}[-5y^2] = -5 \cdot (2y) = -10y

Step 2: Combine the terms:

∂f∂y=3x2+0−10y=3x2−10y\frac{\partial f}{\partial y} = 3x^2 + 0 - 10y = \mathbf{3x^2 - 10y}

Problem 3: Multivariable Linear Sum Derivatives

Consider the linear neuron pre-activation function with two inputs and a bias:

z(w1,w2,b)=w1x1+w2x2+bz(w_1, w_2, b) = w_1 x_1 + w_2 x_2 + b

Given fixed feature measurements x1=3.0x_1 = 3.0 and x2=−4.0x_2 = -4.0:

  • Part A: Compute the analytical partial derivatives ∂z∂w1\frac{\partial z}{\partial w_1}, ∂z∂w2\frac{\partial z}{\partial w_2}, and ∂z∂b\frac{\partial z}{\partial b}.
  • Part B: Evaluate their exact numerical values.
Reveal Solution

Part A: Analytical Partial Derivatives

  • ∂z∂w1=∂∂w1[w1x1+w2x2+b]=x1+0+0=x1\frac{\partial z}{\partial w_1} = \frac{\partial}{\partial w_1}[w_1 x_1 + w_2 x_2 + b] = x_1 + 0 + 0 = \mathbf{x_1}
  • ∂z∂w2=∂∂w2[w1x1+w2x2+b]=0+x2+0=x2\frac{\partial z}{\partial w_2} = \frac{\partial}{\partial w_2}[w_1 x_1 + w_2 x_2 + b] = 0 + x_2 + 0 = \mathbf{x_2}
  • ∂z∂b=∂∂b[w1x1+w2x2+b]=0+0+1=1\frac{\partial z}{\partial b} = \frac{\partial}{\partial b}[w_1 x_1 + w_2 x_2 + b] = 0 + 0 + 1 = \mathbf{1}

Part B: Numerical Evaluation Substituting x1=3.0x_1 = 3.0 and x2=−4.0x_2 = -4.0:

  • ∂z∂w1=+3.0\frac{\partial z}{\partial w_1} = \mathbf{+3.0}
  • ∂z∂w2=−4.0\frac{\partial z}{\partial w_2} = \mathbf{-4.0}
  • ∂z∂b=+1.0\frac{\partial z}{\partial b} = \mathbf{+1.0}

Interpretation: Increasing weight w1w_1 by +0.1+0.1 increases zz by +0.3+0.3. Increasing weight w2w_2 by +0.1+0.1 decreases zz by −0.4-0.4.


Problem 4: Gradient Vector Assembly and Magnitude

Given the scalar loss function:

L(w1,w2)=2w12+3w22−4w1+6w2+5L(w_1, w_2) = 2w_1^2 + 3w_2^2 - 4w_1 + 6w_2 + 5
  • Part A: Compute the general partial derivative formulas ∂L∂w1\frac{\partial L}{\partial w_1} and ∂L∂w2\frac{\partial L}{\partial w_2}.
  • Part B: Evaluate the partial derivatives at the operating point (w1,w2)=(2,−1)(w_1, w_2) = (2, -1).
  • Part C: Assemble the 2D gradient vector ∇L\nabla L at this point and compute its magnitude ∥∇L∥\|\nabla L\|.
Reveal Solution

Part A: General Partial Derivatives

  • ∂L∂w1=∂∂w1[2w12−4w1+constant]=4w1−4\frac{\partial L}{\partial w_1} = \frac{\partial}{\partial w_1}[2w_1^2 - 4w_1 + \text{constant}] = 4w_1 - 4
  • ∂L∂w2=∂∂w2[3w22+6w2+constant]=6w2+6\frac{\partial L}{\partial w_2} = \frac{\partial}{\partial w_2}[3w_2^2 + 6w_2 + \text{constant}] = 6w_2 + 6

Part B: Evaluate at (w1,w2)=(2,−1)(w_1, w_2) = (2, -1)

  • ∂L∂w1∣(2,−1)=4(2)−4=8−4=+4.0\left. \frac{\partial L}{\partial w_1} \right|_{(2, -1)} = 4(2) - 4 = 8 - 4 = \mathbf{+4.0}
  • ∂L∂w2∣(2,−1)=6(−1)+6=−6+6=0.0\left. \frac{\partial L}{\partial w_2} \right|_{(2, -1)} = 6(-1) + 6 = -6 + 6 = \mathbf{0.0}

Part C: Gradient Vector and Magnitude Assemble the column vector:

∇L(2,−1)=[∂L∂w1∂L∂w2]=[+4.00.0]\nabla L(2, -1) = \begin{bmatrix} \frac{\partial L}{\partial w_1} \\ \frac{\partial L}{\partial w_2} \end{bmatrix} = \begin{bmatrix} +4.0 \\ 0.0 \end{bmatrix}

Compute the Euclidean magnitude:

∥∇L∥=(+4.0)2+(0.0)2=16.0=4.0\|\nabla L\| = \sqrt{(+4.0)^2 + (0.0)^2} = \sqrt{16.0} = \mathbf{4.0}

Geometric Conclusion: At coordinate (2,−1)(2, -1), the loss surface is flat along the w2w_2 axis (∂L∂w2=0\frac{\partial L}{\partial w_2} = 0), and slopes steeply uphill along the w1w_1 axis with a slope of +4.0+4.0.


Problem 5: Steepest Descent Step and Error Reduction Verification

Using the loss function and gradient from Problem 4 (L(w1,w2)=2w12+3w22−4w1+6w2+5L(w_1, w_2) = 2w_1^2 + 3w_2^2 - 4w_1 + 6w_2 + 5 at (w1,w2)=(2,−1)(w_1, w_2) = (2, -1)):

  • Part A: Write down the steepest descent vector −∇L-\nabla L.
  • Part B: Perform a parameter update step with learning rate η=0.10\eta = 0.10: wnew=wold−η∇Lw_{\text{new}} = w_{\text{old}} - \eta \nabla L
  • Part C: Evaluate the initial loss L(wold)L(w_{\text{old}}) and the updated loss L(wnew)L(w_{\text{new}}). Verify that the step strictly reduced the loss.
Reveal Solution

Part A: Steepest Descent Vector

−∇L=−[+4.00.0]=[−4.00.0]-\nabla L = -\begin{bmatrix} +4.0 \\ 0.0 \end{bmatrix} = \begin{bmatrix} -4.0 \\ 0.0 \end{bmatrix}

Part B: Compute Updated Parameter Coordinates

wnew=[2.0−1.0]−0.10[4.00.0]=[2.0−0.40−1.0−0.00]=[1.60−1.00]w_{\text{new}} = \begin{bmatrix} 2.0 \\ -1.0 \end{bmatrix} - 0.10 \begin{bmatrix} 4.0 \\ 0.0 \end{bmatrix} = \begin{bmatrix} 2.0 - 0.40 \\ -1.0 - 0.00 \end{bmatrix} = \begin{bmatrix} \mathbf{1.60} \\ \mathbf{-1.00} \end{bmatrix}

Part C: Loss Evaluation and Verification

  • Initial Loss at (2.0,−1.0)(2.0, -1.0): L(2.0,−1.0)=2(2.0)2+3(−1.0)2−4(2.0)+6(−1.0)+5L(2.0, -1.0) = 2(2.0)^2 + 3(-1.0)^2 - 4(2.0) + 6(-1.0) + 5 L(2.0,−1.0)=2(4)+3(1)−8−6+5=8+3−8−6+5=2.00L(2.0, -1.0) = 2(4) + 3(1) - 8 - 6 + 5 = 8 + 3 - 8 - 6 + 5 = \mathbf{2.00}

  • Updated Loss at (1.60,−1.00)(1.60, -1.00): L(1.60,−1.00)=2(1.60)2+3(−1.00)2−4(1.60)+6(−1.00)+5L(1.60, -1.00) = 2(1.60)^2 + 3(-1.00)^2 - 4(1.60) + 6(-1.00) + 5 L(1.60,−1.00)=2(2.56)+3(1)−6.40−6.00+5L(1.60, -1.00) = 2(2.56) + 3(1) - 6.40 - 6.00 + 5 L(1.60,−1.00)=5.12+3.00−6.40−6.00+5.00=0.72L(1.60, -1.00) = 5.12 + 3.00 - 6.40 - 6.00 + 5.00 = \mathbf{0.72}

  • Loss Reduction (ΔL\Delta L): ΔL=0.72−2.00=−1.28\Delta L = 0.72 - 2.00 = \mathbf{-1.28}

Mathematical Proof: Stepping along the negative gradient reduced the loss penalty from 2.002.00 to 0.720.72 (a 64%64\% error reduction in a single step).


Part 2: Applied Scenario: VC Portfolio Risk Optimization

In our venture capital firm, an automated risk model evaluates startup investment criteria across two strategic dimensions:

  • w1w_1: Market Expansion Traction Weight
  • w2w_2: Unit Economics Margin Weight

The firm's historical portfolio risk loss function is given by the anisotropic quadratic error bowl:

L(w1,w2)=(w1−3)2+2(w2−5)2L(w_1, w_2) = (w_1 - 3)^2 + 2(w_2 - 5)^2

The optimal risk-minimizing criteria target coordinates are (w1∗,w2∗)=(3.0,5.0)(w_1^*, w_2^*) = (3.0, 5.0), where portfolio risk loss reaches its absolute minimum of zero (L=0.0L = 0.0).

A new partner at the firm is currently evaluating startup pitches using aggressively over-weighted criteria coordinates:

(w1,w2)=(5.0,8.0)(w_1, w_2) = (5.0, 8.0)


Problem 6: Step-by-Step Multivariable Risk Gradient Audit

  • Part A: Compute the current portfolio risk loss penalty L(5.0,8.0)L(5.0, 8.0).
  • Part B: Derive the analytical partial derivative formulas ∂L∂w1\frac{\partial L}{\partial w_1} and ∂L∂w2\frac{\partial L}{\partial w_2}.
  • Part C: Evaluate the partial derivatives at (w1,w2)=(5.0,8.0)(w_1, w_2) = (5.0, 8.0), assemble the 2D gradient vector ∇L\nabla L, and compute its magnitude ∥∇L∥\|\nabla L\|.
  • Part D (Steepest Descent Step):
    1. Determine the steepest descent direction vector −∇L-\nabla L.
    2. Perform a test parameter update step with step size η=0.25\eta = 0.25: wnew=wold−η∇Lw_{\text{new}} = w_{\text{old}} - \eta \nabla L
    3. Evaluate the new portfolio risk loss L(wnew)L(w_{\text{new}}).
    4. Calculate the total loss reduction ΔL\Delta L and percentage error eliminated.
  • Part E (Investor Synthesis): In 2–3 sentences, explain why the gradient component for Unit Economics (∂L∂w2=12.0\frac{\partial L}{\partial w_2} = 12.0) is three times larger than Market Expansion (∂L∂w1=4.0\frac{\partial L}{\partial w_1} = 4.0), and how multivariable gradient descent systematically prioritizes corrections on the most severely distorted investment criteria.
Reveal Solution

Part A: Current Portfolio Risk Loss

L(5.0,8.0)=(5.0−3)2+2(8.0−5)2=(2.0)2+2(3.0)2=4.0+2(9.0)=4.0+18.0=22.00L(5.0, 8.0) = (5.0 - 3)^2 + 2(8.0 - 5)^2 = (2.0)^2 + 2(3.0)^2 = 4.0 + 2(9.0) = 4.0 + 18.0 = \mathbf{22.00}

The current partner criteria configuration incurs a severe portfolio risk loss score of 22.0022.00.


Part B: Analytical Partial Derivatives

Differentiate L(w1,w2)=(w1−3)2+2(w2−5)2L(w_1, w_2) = (w_1 - 3)^2 + 2(w_2 - 5)^2 term-by-term:

  1. Partial derivative with respect to w1w_1 (freezing w2w_2): ∂L∂w1=∂∂w1[(w1−3)2]+0=2(w1−3)⋅1=2(w1−3)\frac{\partial L}{\partial w_1} = \frac{\partial}{\partial w_1}[(w_1 - 3)^2] + 0 = 2(w_1 - 3) \cdot 1 = \mathbf{2(w_1 - 3)}

  2. Partial derivative with respect to w2w_2 (freezing w1w_1): ∂L∂w2=0+∂∂w2[2(w2−5)2]=2⋅2(w2−5)⋅1=4(w2−5)\frac{\partial L}{\partial w_2} = 0 + \frac{\partial}{\partial w_2}[2(w_2 - 5)^2] = 2 \cdot 2(w_2 - 5) \cdot 1 = \mathbf{4(w_2 - 5)}


Part C: Gradient Vector Assembly and Magnitude at (5.0,8.0)(5.0, 8.0)

Evaluate the partial derivatives at w1=5.0w_1 = 5.0 and w2=8.0w_2 = 8.0:

  • ∂L∂w1∣(5.0,8.0)=2(5.0−3)=2(2.0)=+4.0\left. \frac{\partial L}{\partial w_1} \right|_{(5.0, 8.0)} = 2(5.0 - 3) = 2(2.0) = \mathbf{+4.0}
  • ∂L∂w2∣(5.0,8.0)=4(8.0−5)=4(3.0)=+12.0\left. \frac{\partial L}{\partial w_2} \right|_{(5.0, 8.0)} = 4(8.0 - 5) = 4(3.0) = \mathbf{+12.0}

Assemble the 2D gradient vector:

∇L=[∂L∂w1∂L∂w2]=[+4.0+12.0]\nabla L = \begin{bmatrix} \frac{\partial L}{\partial w_1} \\ \frac{\partial L}{\partial w_2} \end{bmatrix} = \begin{bmatrix} +4.0 \\ +12.0 \end{bmatrix}

Compute the gradient magnitude:

∥∇L∥=(4.0)2+(12.0)2=16.0+144.0=160.0=410≈12.649\|\nabla L\| = \sqrt{(4.0)^2 + (12.0)^2} = \sqrt{16.0 + 144.0} = \sqrt{160.0} = 4\sqrt{10} \approx \mathbf{12.649}

Part D: Steepest Descent Step and Loss Reduction

  1. Steepest Descent Direction Vector: −∇L=[−4.0−12.0]-\nabla L = \begin{bmatrix} -4.0 \\ -12.0 \end{bmatrix}

  2. Parameter Update with η=0.25\eta = 0.25: wnew=[5.08.0]−0.25[4.012.0]=[5.0−1.08.0−3.0]=[4.05.0]w_{\text{new}} = \begin{bmatrix} 5.0 \\ 8.0 \end{bmatrix} - 0.25 \begin{bmatrix} 4.0 \\ 12.0 \end{bmatrix} = \begin{bmatrix} 5.0 - 1.0 \\ 8.0 - 3.0 \end{bmatrix} = \begin{bmatrix} \mathbf{4.0} \\ \mathbf{5.0} \end{bmatrix}

  3. Evaluate Updated Loss L(4.0,5.0)L(4.0, 5.0): L(4.0,5.0)=(4.0−3)2+2(5.0−5)2=(1.0)2+2(0.0)2=1.0+0.0=1.00L(4.0, 5.0) = (4.0 - 3)^2 + 2(5.0 - 5)^2 = (1.0)^2 + 2(0.0)^2 = 1.0 + 0.0 = \mathbf{1.00}

  4. Loss Reduction and Efficiency Audit: ΔL=Lnew−Lold=1.00−22.00=−21.00\Delta L = L_{\text{new}} - L_{\text{old}} = 1.00 - 22.00 = \mathbf{-21.00} Percentage Error Reduced=21.0022.00×100%=95.45%\text{Percentage Error Reduced} = \frac{21.00}{22.00} \times 100\% = \mathbf{95.45\%}

A single gradient descent step reduced the portfolio risk loss by 95.45%95.45\%, pulling Unit Economics (w2w_2) directly onto its target of 5.05.0 and Market Expansion (w1w_1) from 5.05.0 down to 4.04.0.


Part E: Investor Synthesis

The Unit Economics gradient (∂L∂w2=12.0\frac{\partial L}{\partial w_2} = 12.0) is three times larger than the Market Expansion gradient (∂L∂w1=4.0\frac{\partial L}{\partial w_1} = 4.0) for two compound reasons: the partner was further away from the target along w2w_2 (Δw2=3.0\Delta w_2 = 3.0 vs Δw1=2.0\Delta w_1 = 2.0), and the loss surface penalizes Unit Economics errors twice as heavily (curvature factor of 22).

Multivariable gradient vectors automatically account for both distance and surface curvature, directing the largest corrective updates to the specific parameters that contribute most heavily to risk.

Previous
Direction of Steepest Ascent and Descent