The Backward Pass
Backward pass mechanics in neural networks: loss functions, single-variable derivatives, multivariable gradient vectors, and chain rule backpropagation.
01 • Prediction Error & Loss Functions
Grounding Prediction Error in Reality
Quantify neural network prediction errors using Mean Squared Error, analyze parabolic loss curves, and derive loss sensitivities to predictions.
The Mean Squared Error Loss Function
Formulate the Mean Squared Error loss function to convert prediction mistakes into positive, differentiable scalar penalties for neural networks.
Parabolic Loss Surfaces and Curvature
Analyze the parabolic geometry of squared error curves to visualize error minimums, steep loss walls, and symmetric penalties around target values.
Output Loss Derivatives and Gradients
Compute the derivative of the loss function with respect to predicted outputs, establishing the first mathematical link in error attribution.
Prediction Error & Loss Functions In Practice
Master raw error calculation, Mean Squared Error formulas, parabolic curvature geometry, and output loss derivatives through manual calculations.
02 • Derivatives and Sensitivity
The Rate of Change and Tangent Slopes
Master single-variable calculus foundations, transition from intuitive sensitivity to formal derivatives, and apply power rules to tangent slopes.
Limits and the Formal Derivative Definition
Define the formal mathematical derivative as the limit of average rates of change as the step interval shrinks toward an infinitesimal nudge.
The Training Wheels Transition in Calculus
Transition from intuitive sensitivity terminology to formal derivative notation, analyzing positive, negative, and zero slope conditions in depth.
Elementary Power and Sum Derivative Rules
Apply foundational power, constant, and sum rules of differential calculus to compute exact algebraic derivatives for single-variable functions.
Derivatives and Sensitivity In Practice
Master secant line limits, algebraic derivative rules, stationary slope conditions, and valuation sensitivity rates through manual calculations.
03 • Partial Derivatives and Gradients
Multivariable Systems and Freezing
Isolate causal parameter effects in multivariable systems, calculate weight partial derivatives, and assemble gradient vectors for error descent.
Calculating Partial Derivatives of Weights
Compute partial derivatives for individual parameter weights under Mean Squared Error loss by treating all competing weights as static constants.
Assembling the Multivariable Gradient Vector
Assemble individual partial derivatives into a unified gradient vector matching the exact dimensionality of the neural network parameter space.
Direction of Steepest Ascent and Descent
Analyze the geometric orientation of gradient vectors to establish why the negative gradient points along the path of steepest error reduction.
Partial Derivatives & Gradient Vectors In Practice
Master multivariable parameter isolation, partial derivative calculations, gradient vector assembly, and steepest descent trajectories through manual calculations.
04 • The Chain Rule and Backpropagation
The Composite Function Chain Rule
Apply composite chain rules to attribute error through output and hidden layers, computing backward gradients and toy parameter updates on paper.
Output Layer Error Attribution Math
Decompose output layer gradients into loss sensitivity, activation slope, and input terms to compute exact credit attribution for output weights.
Hidden Layer Error Backpropagation
Propagate error signals backward through hidden layers by multiplying downstream deltas by transposed weights and intermediate activation slopes.
Toy Parameter Update Step on Paper
Perform a single conceptual parameter update step by scaling computed gradients with a step factor to demonstrate how weights adjust on paper.
The Chain Rule & Backpropagation In Practice
Master composite chain rules, output layer error attribution, hidden backpropagation deltas, and single-step parameter updates through manual calculations.