Multivariable Systems and Freezing hero
Lesson 1Partial Derivatives and Gradients

Multivariable Systems and Freezing

Isolate causal parameter effects in multivariable systems, calculate weight partial derivatives, and assemble gradient vectors for error descent.

In Topic 1, we measured prediction error on the screenplay for 'Die Hard in Space' using Mean Squared Error:

L=(y−y^)2=(0−0.95)2=0.9025L = (y - \hat{y})^2 = (0 - 0.95)^2 = \mathbf{0.9025}

In Topic 2, we established the mathematical machinery of single-variable calculus. We proved that the derivative dydx\frac{dy}{dx} measures the instantaneous rate of change of an output when a single input is nudged by an infinitesimal amount.

[ Topic 1: Loss Functions ] ──► [ Topic 2: Derivatives ] ──► [ Topic 3: Gradients ] ──► [ Topic 4: Backpropagation ]
      L = (y - y_hat)^2               dy/dx & f'(x)                 grad_w L = dL/dw             Chain Rule Across Layers

However, neural networks never operate on a single parameter in isolation.

When our network evaluated 'Die Hard in Space', it combined multiple genre feature measurements—Action (x1x_1), Romance (x2x_2), Comedy (x3x_3), and Sci-Fi (x4x_4)—with corresponding parameter weights (w1,w2,w3,w4w_1, w_2, w_3, w_4) and a bias (bb) into an affine linear pre-activation:

z=w1x1+w2x2+w3x3+w4x4+bz = w_1 x_1 + w_2 x_2 + w_3 x_3 + w_4 x_4 + b

Because the final loss penalty LL depends simultaneously on all of these parameters, changing multiple weights at once makes it impossible to know which parameter helped and which hurt.

To solve this credit assignment challenge, multivariable calculus provides two essential tools:

  1. Partial Derivatives (∂L∂wi\frac{\partial L}{\partial w_i}): Isolate the causal effect of an individual parameter by freezing all other parameters as static constants.
  2. The Gradient Vector (∇L\nabla L): Bundle all individual partial derivatives into a single directional vector that points along the path of steepest error change in parameter space.


Isolate individual parameters in multivariable systems by holding other inputs constant to evaluate isolated causal effects on network outcomes.

In single-variable calculus, a function takes one input number and produces one output number:

y=f(x)y = f(x)

In deep learning, loss functions depend on hundreds, thousands, or billions of interconnected parameters:

L=f(w1,w2,w3,…,wn,b)L = f(w_1, w_2, w_3, \dots, w_n, b)

The Challenge of Simultaneous Change: The Causal Credit Problem

Imagine adjusting an audio mixing console with four distinct frequency sliders: Bass, Low-Mid, High-Mid, and Treble.

If you push the Bass slider up by +3 dB+3\text{ dB}, pull the Low-Mid slider down by −2 dB-2\text{ dB}, push the High-Mid slider up by +4 dB+4\text{ dB}, and pull Treble down by −1 dB-1\text{ dB} all in the same second, the overall acoustic balance changes.

If the sound becomes harsh and distorted, which specific slider caused the distortion?

Simultaneous Adjustments:
  Bass     (+3 dB) ──┐
  Low-Mid  (-2 dB) ──┼──► [ Total Audio Mix ] ──► Harsh Distortion Occurred
  High-Mid (+4 dB) ──┤                            (Which slider is responsible?)
  Treble   (-1 dB) ──┘

Because four dials moved at the same time, their causal effects are hopelessly entangled. You cannot determine whether the High-Mid boost was entirely responsible, or if the Bass boost compounded the issue.

This is the Causal Credit Problem. If a neural network updates all its parameter weights simultaneously without knowing each parameter's individual sensitivity, it cannot determine which weights reduced the error and which made it worse.


The Parameter Isolation Principle (Freezing)

The scientific method solves simultaneous confounding through controlled isolation: to determine the effect of one variable, hold all other variables constant.

Multivariable calculus applies this exact isolation principle. To evaluate how sensitive the loss LL is to a specific weight w1w_1:

  1. Freeze all competing weights (w2,w3,…,wnw_2, w_3, \dots, w_n) and bias bb at their current numerical values.
  2. Treat those frozen parameters as static, unmoving constants (identical to fixed numbers like 55 or −12-12).
  3. Nudge only the single target parameter w1w_1 by an infinitesimal amount Δw1\Delta w_1.
  4. Measure the resulting change in loss ΔL\Delta L.
Parameter Isolation (Freezing):
  w_1 (Action)   ──► Nudged by Δw_1  ──► Active Variable Under Study
  w_2 (Romance)  ──► FROZEN (Constant)
  w_3 (Comedy)   ──► FROZEN (Constant)
  w_4 (Sci-Fi)   ──► FROZEN (Constant)
  b   (Bias)     ──► FROZEN (Constant)

By freezing all other parameters, the multivariable function collapses into a simple 1D single-variable slice along the w1w_1 axis. We can then apply the standard single-variable calculus rules mastered in Topic 2.


Dual-Track Bridge: Isolating Mechanisms

Track 1 (Underlying Mechanism): Mountain Terrain in 3D Space

Imagine standing on the side of a mountain at coordinates (x0,y0)(x_0, y_0) with elevation z=f(x,y)z = f(x, y), where xx represents your North-South coordinate and yy represents your East-West coordinate.

                  North (x)
                     ▲
                     │     Slope along North-South: Freeze y = y_0
                     │     ∂z/∂x = Rate of elevation change moving North
                     │
                     ┼──────────────► East (y)
                   (x_0, y_0)   Slope along East-West: Freeze x = x_0
                                ∂z/∂y = Rate of elevation change moving East
  • To measure the slope in the North-South direction, you freeze your East-West coordinate (y=y0y = y_0) and take a small step North. The rate of elevation change per meter walked North is the partial rate of change with respect to xx.
  • To measure the slope in the East-West direction, you freeze your North-South coordinate (x=x0x = x_0) and take a small step East. The rate of elevation change per meter walked East is the partial rate of change with respect to yy.

Track 2 (Applied Concept): Audience Genre Weights on 'Die Hard in Space'

In our Netflix greenlight predictor, the pre-activation sum is:

z(w1,w2,w3,w4)=w1x1+w2x2+w3x3+w4x4+bz(w_1, w_2, w_3, w_4) = w_1 x_1 + w_2 x_2 + w_3 x_3 + w_4 x_4 + b

For 'Die Hard in Space', the feature measurements are locked: Action x1=0.90x_1 = 0.90, Romance x2=0.10x_2 = 0.10, Comedy x3=0.20x_3 = 0.20, and Sci-Fi x4=0.80x_4 = 0.80.

To isolate how the Action weight (w1w_1) affects the prediction, we freeze Romance (w2=0.10w_2 = 0.10), Comedy (w3=0.20w_3 = 0.20), Sci-Fi (w4=0.70w_4 = 0.70), and bias (b=0.05b = 0.05). The linear sum simplifies to a single-variable function of w1w_1:

z(w1)=w1(0.90)+(0.10)(0.10)+(0.20)(0.20)+(0.70)(0.80)+0.05z(w_1) = w_1(0.90) + (0.10)(0.10) + (0.20)(0.20) + (0.70)(0.80) + 0.05 z(w1)=0.90w1+[0.01+0.04+0.56+0.05]=0.90w1+0.66z(w_1) = 0.90 w_1 + [0.01 + 0.04 + 0.56 + 0.05] = 0.90 w_1 + 0.66

All frozen terms collapse into a single constant (+0.66+0.66). The rate of change of zz with respect to w1w_1 is simply the feature measurement 0.900.90.


Previous
Derivatives and Sensitivity In Practice