The Forward Pass
Forward data flow in neural networks: feature vectors, dot products, baseline biases, activation functions, and multi-layer dense matrix layers.
01 • Vectors & Dot Products (Linear Algebra Part 1)
Feature Vectors
Ordered numeric lists and arrays that establish mathematical identity profiles for real-world attributes in neural network layers.
Vector Addition and Scaling
Position-wise vector addition and scalar multiplication that combine feature profiles, scale attribute values, and combine multidimensional data.
Dot Products and Weights
Pairwise feature-weight multiplication and summation that accumulate attribute influences into single scalar decision scores in neural network layers.
From Vectors to Neural Networks
The bridge from raw dot products to artificial neurons, exposing uncalibrated score scales and motivating baseline bias offsets and activation functions.
Vectors and Dot Products In Practice
Manual problem sets covering vector construction, position-wise arithmetic, and dot products with weights in abstract and applied decision settings.
02 • The Artificial Neuron & Activations
Linear Step and Baseline Bias
Baseline bias offsets that establish independent starting thresholds, completing the linear step and shifting dot products along the decision threshold.
Sigmoid Activation Function
The S-shaped Sigmoid activation function that squashes unbounded linear scores into standardized probabilities between zero and one.
Artificial Neuron Architecture
The unified computational pipeline combining input features, influence weights, baseline bias offsets, and non-linear activation functions.
Artificial Neuron In Practice
Manual step-by-step calculations of linear scores, Sigmoid probabilities, and baseline bias offsets to evaluate decision thresholds.
03 • Matrices & Layer Width (Linear Algebra Part 2)
Matrix Dimensions and Parallel Vectors
Stack individual neuron weight vectors into two-dimensional matrices to evaluate multiple outcomes on the same input vector at once.
Matrix-Vector Multiplication
Matrix-vector multiplication executed as parallel row dot products to transform input feature coordinates into multi-output decision vectors.
Layer Width and Parallel Decisions
Define layer width by the count of parallel neurons evaluating a shared input vector, and establish independent decision hurdles with bias vectors.
The Affine Layer Transformation Math
Affine transformations and coordinate-wise activations that assemble matrix products and bias vectors into the complete forward pass of a dense layer.
Matrices and Layer Width In Practice
Master matrix-vector multiplication, layer width scaling, affine layer transformations, and parallel multi-output evaluation through manual calculation.
04 • Multi-Layer Neural Networks
Limitations of Single-Layer Networks
Explore why single-layer linear networks cannot learn compound feature interactions or solve non-linearly separable classification problems.
Hidden Layers and Network Geometry
Insert intermediate hidden layers between inputs and outputs to create hierarchical geometric transformations across multi-layer perceptrons.
Representation Learning Foundations
Trace how successive hidden layers automatically transform low-level input features into high-level abstract representations inside the network.
The Rectified Linear Unit Activation
Apply the piecewise linear ReLU function to introduce non-linearity while preserving identity gradients for positive activation values in neurons.
The Complete MLP Forward Pass Math
Formulate the end-to-end forward pass equation chaining matrix transformations, biases, and activations from input vector to final predictions.
Multi-Layer Networks In Practice
Master multi-layer forward passes, hidden representations, composite activations, and dimension tracking through complete manual calculations.