Anisotropic Loss Bowl: L(w1,w2)=4(w1−2)2+(w2−3)2 (steeper curvature along w1 than w2).
Flop Error Coordinate: High-error operating point at (w1,w2)=(5.0,4.0) demonstrating large gradient magnitude.
Optimal Valley Minimum: Global minimum at (w1∗,w2∗)=(2.0,3.0) where ∇L=[00] and ∥∇L∥=0.
Guided Challenge Cue:
"Challenge: Lock w2 and adjust w1 to find the 1D slice minimum, then release w2 and observe how the true 2D gradient points along a combined descent trajectory perpendicular to the contour lines."
Part 1: The Underlying Mechanics Drill
Problem 1: Partial Derivative with Respect to x
Given the multivariable polynomial function:
f(x,y)=3x2y+4x−5y2
Compute the partial derivative ∂x∂f by treating y as a constant.
Reveal Solution
Step 1: Differentiate each term with respect to x, treating y as a constant:
Term 1 (3x2y): The factor 3y acts as a constant multiplier on x2:
∂x∂[3y⋅x2]=3y⋅(2x)=6xy
Term 2 (4x): Linear multiplier rule:
∂x∂[4x]=4
Term 3 (−5y2): Contains no x variable. Since y is constant, −5y2 is constant:
∂x∂[−5y2]=0
Step 2: Combine the terms:
∂x∂f=6xy+4+0=6xy+4
Problem 2: Partial Derivative with Respect to y
Using the same function:
f(x,y)=3x2y+4x−5y2
Compute the partial derivative ∂y∂f by treating x as a constant.
Reveal Solution
Step 1: Differentiate each term with respect to y, treating x as a constant:
Term 1 (3x2y): The factor 3x2 acts as a constant multiplier on y:
∂y∂[3x2⋅y]=3x2⋅(1)=3x2
Term 2 (4x): Contains no y variable. Since x is constant, 4x is constant:
∂y∂[4x]=0
Term 3 (−5y2): Standard power rule on y:
∂y∂[−5y2]=−5⋅(2y)=−10y
Step 2: Combine the terms:
∂y∂f=3x2+0−10y=3x2−10y
Problem 3: Multivariable Linear Sum Derivatives
Consider the linear neuron pre-activation function with two inputs and a bias:
z(w1,w2,b)=w1x1+w2x2+b
Given fixed feature measurements x1=3.0 and x2=−4.0:
Part A: Compute the analytical partial derivatives ∂w1∂z, ∂w2∂z, and ∂b∂z.
Part B: Evaluate their exact numerical values.
Reveal Solution
Part A: Analytical Partial Derivatives
∂w1∂z=∂w1∂[w1x1+w2x2+b]=x1+0+0=x1
∂w2∂z=∂w2∂[w1x1+w2x2+b]=0+x2+0=x2
∂b∂z=∂b∂[w1x1+w2x2+b]=0+0+1=1
Part B: Numerical Evaluation
Substituting x1=3.0 and x2=−4.0:
∂w1∂z=+3.0
∂w2∂z=−4.0
∂b∂z=+1.0
Interpretation: Increasing weight w1 by +0.1 increases z by +0.3. Increasing weight w2 by +0.1 decreases z by −0.4.
Problem 4: Gradient Vector Assembly and Magnitude
Given the scalar loss function:
L(w1,w2)=2w12+3w22−4w1+6w2+5
Part A: Compute the general partial derivative formulas ∂w1∂L and ∂w2∂L.
Part B: Evaluate the partial derivatives at the operating point (w1,w2)=(2,−1).
Part C: Assemble the 2D gradient vector ∇L at this point and compute its magnitude ∥∇L∥.
Reveal Solution
Part A: General Partial Derivatives
∂w1∂L=∂w1∂[2w12−4w1+constant]=4w1−4
∂w2∂L=∂w2∂[3w22+6w2+constant]=6w2+6
Part B: Evaluate at (w1,w2)=(2,−1)
∂w1∂L(2,−1)=4(2)−4=8−4=+4.0
∂w2∂L(2,−1)=6(−1)+6=−6+6=0.0
Part C: Gradient Vector and Magnitude
Assemble the column vector:
∇L(2,−1)=[∂w1∂L∂w2∂L]=[+4.00.0]
Compute the Euclidean magnitude:
∥∇L∥=(+4.0)2+(0.0)2=16.0=4.0
Geometric Conclusion: At coordinate (2,−1), the loss surface is flat along the w2 axis (∂w2∂L=0), and slopes steeply uphill along the w1 axis with a slope of +4.0.
Problem 5: Steepest Descent Step and Error Reduction Verification
Using the loss function and gradient from Problem 4 (L(w1,w2)=2w12+3w22−4w1+6w2+5 at (w1,w2)=(2,−1)):
Part A: Write down the steepest descent vector −∇L.
Part B: Perform a parameter update step with learning rate η=0.10:
wnew=wold−η∇L
Part C: Evaluate the initial loss L(wold) and the updated loss L(wnew). Verify that the step strictly reduced the loss.
Initial Loss at (2.0,−1.0):L(2.0,−1.0)=2(2.0)2+3(−1.0)2−4(2.0)+6(−1.0)+5L(2.0,−1.0)=2(4)+3(1)−8−6+5=8+3−8−6+5=2.00
Updated Loss at (1.60,−1.00):L(1.60,−1.00)=2(1.60)2+3(−1.00)2−4(1.60)+6(−1.00)+5L(1.60,−1.00)=2(2.56)+3(1)−6.40−6.00+5L(1.60,−1.00)=5.12+3.00−6.40−6.00+5.00=0.72
Loss Reduction (ΔL):ΔL=0.72−2.00=−1.28
Mathematical Proof: Stepping along the negative gradient reduced the loss penalty from 2.00 to 0.72 (a 64% error reduction in a single step).
Part 2: Applied Scenario: VC Portfolio Risk Optimization
In our venture capital firm, an automated risk model evaluates startup investment criteria across two strategic dimensions:
w1: Market Expansion Traction Weight
w2: Unit Economics Margin Weight
The firm's historical portfolio risk loss function is given by the anisotropic quadratic error bowl:
L(w1,w2)=(w1−3)2+2(w2−5)2
The optimal risk-minimizing criteria target coordinates are (w1∗,w2∗)=(3.0,5.0), where portfolio risk loss reaches its absolute minimum of zero (L=0.0).
A new partner at the firm is currently evaluating startup pitches using aggressively over-weighted criteria coordinates:
(w1,w2)=(5.0,8.0)
Problem 6: Step-by-Step Multivariable Risk Gradient Audit
Part A: Compute the current portfolio risk loss penalty L(5.0,8.0).
Part B: Derive the analytical partial derivative formulas ∂w1∂L and ∂w2∂L.
Part C: Evaluate the partial derivatives at (w1,w2)=(5.0,8.0), assemble the 2D gradient vector ∇L, and compute its magnitude ∥∇L∥.
Part D (Steepest Descent Step):
Determine the steepest descent direction vector −∇L.
Perform a test parameter update step with step size η=0.25:
wnew=wold−η∇L
Evaluate the new portfolio risk loss L(wnew).
Calculate the total loss reduction ΔL and percentage error eliminated.
Part E (Investor Synthesis): In 2–3 sentences, explain why the gradient component for Unit Economics (∂w2∂L=12.0) is three times larger than Market Expansion (∂w1∂L=4.0), and how multivariable gradient descent systematically prioritizes corrections on the most severely distorted investment criteria.
Steepest Descent Direction Vector:−∇L=[−4.0−12.0]
Parameter Update with η=0.25:wnew=[5.08.0]−0.25[4.012.0]=[5.0−1.08.0−3.0]=[4.05.0]
Evaluate Updated Loss L(4.0,5.0):L(4.0,5.0)=(4.0−3)2+2(5.0−5)2=(1.0)2+2(0.0)2=1.0+0.0=1.00
Loss Reduction and Efficiency Audit:ΔL=Lnew−Lold=1.00−22.00=−21.00Percentage Error Reduced=22.0021.00×100%=95.45%
A single gradient descent step reduced the portfolio risk loss by 95.45%, pulling Unit Economics (w2) directly onto its target of 5.0 and Market Expansion (w1) from 5.0 down to 4.0.
Part E: Investor Synthesis
The Unit Economics gradient (∂w2∂L=12.0) is three times larger than the Market Expansion gradient (∂w1∂L=4.0) for two compound reasons: the partner was further away from the target along w2 (Δw2=3.0 vs Δw1=2.0), and the loss surface penalizes Unit Economics errors twice as heavily (curvature factor of 2).
Multivariable gradient vectors automatically account for both distance and surface curvature, directing the largest corrective updates to the specific parameters that contribute most heavily to risk.