
AI Foundations: From Math to Neural Networks
A first-principles mathematical breakdown of vanilla multi-layer neural networks. Covers feature vectors, dense layer transformations, loss functions, multivariable gradients, and backpropagation traces without external framework abstractions.
Module 1The Forward Pass
Forward data flow in neural networks: feature vectors, dot products, baseline biases, activation functions, and multi-layer dense matrix layers.
The Forward Pass
Forward data flow in neural networks: feature vectors, dot products, baseline biases, activation functions, and multi-layer dense matrix layers.
01 • Vectors & Dot Products (Linear Algebra Part 1)
Feature Vectors
Ordered numeric lists and arrays that establish mathematical identity profiles for real-world attributes in neural network layers.
Vector Addition and Scaling
Position-wise vector addition and scalar multiplication that combine feature profiles, scale attribute values, and combine multidimensional data.
Dot Products and Weights
Pairwise feature-weight multiplication and summation that accumulate attribute influences into single scalar decision scores in neural network layers.
From Vectors to Neural Networks
The bridge from raw dot products to artificial neurons, exposing uncalibrated score scales and motivating baseline bias offsets and activation functions.
Vectors and Dot Products In Practice
Manual problem sets covering vector construction, position-wise arithmetic, and dot products with weights in abstract and applied decision settings.
02 • The Artificial Neuron & Activations
Linear Step and Baseline Bias
Baseline bias offsets that establish independent starting thresholds, completing the linear step and shifting dot products along the decision threshold.
Sigmoid Activation Function
The S-shaped Sigmoid activation function that squashes unbounded linear scores into standardized probabilities between zero and one.
Artificial Neuron Architecture
The unified computational pipeline combining input features, influence weights, baseline bias offsets, and non-linear activation functions.
Artificial Neuron In Practice
Manual step-by-step calculations of linear scores, Sigmoid probabilities, and baseline bias offsets to evaluate decision thresholds.
03 • Matrices & Layer Width (Linear Algebra Part 2)
Matrix Dimensions and Parallel Vectors
Stack individual neuron weight vectors into two-dimensional matrices to evaluate multiple outcomes on the same input vector at once.
Matrix-Vector Multiplication
Matrix-vector multiplication executed as parallel row dot products to transform input feature coordinates into multi-output decision vectors.
Layer Width and Parallel Decisions
Define layer width by the count of parallel neurons evaluating a shared input vector, and establish independent decision hurdles with bias vectors.
The Affine Layer Transformation Math
Affine transformations and coordinate-wise activations that assemble matrix products and bias vectors into the complete forward pass of a dense layer.
Matrices and Layer Width In Practice
Master matrix-vector multiplication, layer width scaling, affine layer transformations, and parallel multi-output evaluation through manual calculation.
04 • Multi-Layer Neural Networks
Limitations of Single-Layer Networks
Explore why single-layer linear networks cannot learn compound feature interactions or solve non-linearly separable classification problems.
Hidden Layers and Network Geometry
Insert intermediate hidden layers between inputs and outputs to create hierarchical geometric transformations across multi-layer perceptrons.
Representation Learning Foundations
Trace how successive hidden layers automatically transform low-level input features into high-level abstract representations inside the network.
The Rectified Linear Unit Activation
Apply the piecewise linear ReLU function to introduce non-linearity while preserving identity gradients for positive activation values in neurons.
The Complete MLP Forward Pass Math
Formulate the end-to-end forward pass equation chaining matrix transformations, biases, and activations from input vector to final predictions.
Multi-Layer Networks In Practice
Master multi-layer forward passes, hidden representations, composite activations, and dimension tracking through complete manual calculations.
Module 2The Backward Pass
Backward pass mechanics in neural networks: loss functions, single-variable derivatives, multivariable gradient vectors, and chain rule backpropagation.
The Backward Pass
Backward pass mechanics in neural networks: loss functions, single-variable derivatives, multivariable gradient vectors, and chain rule backpropagation.
01 • Prediction Error & Loss Functions
Grounding Prediction Error in Reality
Quantify neural network prediction errors using Mean Squared Error, analyze parabolic loss curves, and derive loss sensitivities to predictions.
The Mean Squared Error Loss Function
Formulate the Mean Squared Error loss function to convert prediction mistakes into positive, differentiable scalar penalties for neural networks.
Parabolic Loss Surfaces and Curvature
Analyze the parabolic geometry of squared error curves to visualize error minimums, steep loss walls, and symmetric penalties around target values.
Output Loss Derivatives and Gradients
Compute the derivative of the loss function with respect to predicted outputs, establishing the first mathematical link in error attribution.
Prediction Error & Loss Functions In Practice
Master raw error calculation, Mean Squared Error formulas, parabolic curvature geometry, and output loss derivatives through manual calculations.
02 • Derivatives and Sensitivity
The Rate of Change and Tangent Slopes
Master single-variable calculus foundations, transition from intuitive sensitivity to formal derivatives, and apply power rules to tangent slopes.
Limits and the Formal Derivative Definition
Define the formal mathematical derivative as the limit of average rates of change as the step interval shrinks toward an infinitesimal nudge.
The Training Wheels Transition in Calculus
Transition from intuitive sensitivity terminology to formal derivative notation, analyzing positive, negative, and zero slope conditions in depth.
Elementary Power and Sum Derivative Rules
Apply foundational power, constant, and sum rules of differential calculus to compute exact algebraic derivatives for single-variable functions.
Derivatives and Sensitivity In Practice
Master secant line limits, algebraic derivative rules, stationary slope conditions, and valuation sensitivity rates through manual calculations.
03 • Partial Derivatives and Gradients
Multivariable Systems and Freezing
Isolate causal parameter effects in multivariable systems, calculate weight partial derivatives, and assemble gradient vectors for error descent.
Calculating Partial Derivatives of Weights
Compute partial derivatives for individual parameter weights under Mean Squared Error loss by treating all competing weights as static constants.
Assembling the Multivariable Gradient Vector
Assemble individual partial derivatives into a unified gradient vector matching the exact dimensionality of the neural network parameter space.
Direction of Steepest Ascent and Descent
Analyze the geometric orientation of gradient vectors to establish why the negative gradient points along the path of steepest error reduction.
Partial Derivatives & Gradient Vectors In Practice
Master multivariable parameter isolation, partial derivative calculations, gradient vector assembly, and steepest descent trajectories through manual calculations.
04 • The Chain Rule and Backpropagation
The Composite Function Chain Rule
Apply composite chain rules to attribute error through output and hidden layers, computing backward gradients and toy parameter updates on paper.
Output Layer Error Attribution Math
Decompose output layer gradients into loss sensitivity, activation slope, and input terms to compute exact credit attribution for output weights.
Hidden Layer Error Backpropagation
Propagate error signals backward through hidden layers by multiplying downstream deltas by transposed weights and intermediate activation slopes.
Toy Parameter Update Step on Paper
Perform a single conceptual parameter update step by scaling computed gradients with a step factor to demonstrate how weights adjust on paper.
The Chain Rule & Backpropagation In Practice
Master composite chain rules, output layer error attribution, hidden backpropagation deltas, and single-step parameter updates through manual calculations.
Module 3Synthesis & Interpretability
Synthesis of the complete forward-backward computational graph, end-to-end mathematical compile traces, and mechanistic interpretability foundations.
Synthesis & Interpretability
Synthesis of the complete forward-backward computational graph, end-to-end mathematical compile traces, and mechanistic interpretability foundations.
01 • The Complete Neural Data Flow Graph
The Forward-Backward Computational Graph
Unify the forward pass, loss calculation, and backward backpropagation into a single directed acyclic graph and verify the closed-loop trace.
The Closed Learning Cycle on Paper
Execute a complete forward pass, loss calculation, backpropagation cycle, and second forward pass on paper to prove error reduction.
The Mathematical Compile Trace Audit
Perform an exhaustive compile trace verifying that every forward tensor and backward gradient calculation relies strictly on course prerequisites.
Operating the Grand Capstone Engine
Operate the Bycroft-grade interactive macro-system to inspect real-time tensor registers and verify hand-calculated values at 60fps.
The Complete Neural Data Flow In Practice
Master complete end-to-end forward and backward passes, loss calculations, gradient attributions, and parameter updates through manual calculations.
02 • Interpretability of Neural Networks
The True Nature of the "Black Box"
Examine the boundary of transparent arithmetic and latent representations, explore mechanistic interpretability, and bridge vanilla MLPs to Transformers.
Interpretable Weights vs. Latent Geometry
Contrast human-interpretable feature weights against high-dimensional latent coordinate geometry inside multi-layer representations.
Mechanistic Interpretability Foundations
Analyze modern scientific methods for probing, ablating, and reverse-engineering the semantic roles of hidden neurons in deep networks.
From Vanilla MLPs to Modern Transformers
Synthesize core vanilla MLP principles and establish the structural bridge to token embeddings, self-attention, and large language models.
Neural Interpretability In Practice
Master latent representation geometry, linear separability, feature ablation, and Transformer MLP sublayer calculations through manual hand traces.