Lesson 4Matrices & Layer Width (Linear Algebra Part 2)
The Affine Layer Transformation Math
Affine transformations and coordinate-wise activations that assemble matrix products and bias vectors into the complete forward pass of a dense layer.
Across the first three lessons of this topic, we constructed the individual mathematical components of a multi-neuron layer:
In Lesson 1, we stacked individual weight vectors into a 2D weight matrix W to evaluate multiple outcomes simultaneously.
In Lesson 2, we multiplied that matrix by an input feature vector x to execute parallel row dot products in a single operation.
In Lesson 3, we introduced the bias vector b, establishing an independent baseline hurdle for each parallel neuron.
Now, we connect these linear components with non-linear activation functions to form the complete mathematical forward pass of a single dense layer. Every fully connected layer in deep learning—from simple binary classifiers to multi-billion-parameter foundation models—computes its output using this exact sequence.
The Complete Dense Layer Pipeline
The forward pass of a single neural network layer consists of two sequential steps: an Affine Linear Transformation followed by a Coordinate-wise Non-Linear Activation.
Step1(AffineLinearTransformation):z=Wx+b
Step2(Coordinate-wiseActivation):a=σ(z)
Why is Step 1 formally called an affine transformation rather than simply a linear transformation?
In pure linear algebra, matrix multiplication (Wx) is strictly linear because it maps zero to zero: if the incoming feature vector contains all zeros (x=0), the matrix product is strictly zero (W0=0). A linear transformation can rotate, scale, or shear a coordinate space, but it can never shift the origin. Adding the bias vector (+b) translates—or shifts—the coordinate space away from the origin (W0+b=b). A linear transformation combined with a translation is mathematically defined as an affine transformation. In deep learning, this affine transformation yields the pre-activation vector z.
Step 2 then applies an activation function σ coordinate-wise to each entry in z. Without this non-linear step, multiple stacked layers would collapse into a single flat linear equation (W2(W1x+b1)+b2=(W2W1)x+(W2b1+b2)). In our standalone studio classifier, the Sigmoid function squashes each unbounded linear score zi into a calibrated decision probability between 0 and 1 (0% and 100%).
Tracing the complete pipeline from input to output:
Matrix Multiplication (Wx): Multiplies the m×n weight matrix by the n×1 input feature vector, executing m parallel row dot products to yield an intermediate m×1 vector.
Vector Addition (+b): Adds the m×1 bias vector element-by-element to the dot products, shifting each channel by its independent baseline hurdle to yield the pre-activation vector z∈Rm.
Coordinate-wise Activation (σ(z)): Applies the activation function independently to each coordinate zi, squashing each linear sum into the final layer activation vector a∈Rm.
Numerical Walk-Through: Script Evaluation
Let's trace how our 3-neuron studio layer evaluates incoming scripts using the parameters established in Lessons 1–3:
aDie Hard=0.9990.0290.998(99.9% Hit Probability)(2.9% Oscar Probability)(99.8% Franchise Probability)
🎬 Script 2: 'The Notebook 2'
To see how the exact same layer weights and biases evaluate an entirely different genre profile, consider 'The Notebook 2' carried forward from Topics 1 and 2:
aNotebook 2=0.0290.9990.047(2.9% Hit Probability)(99.9% Oscar Probability)(4.7% Franchise Probability)
Notice the outcome: the single weight matrix W and bias vector b cleanly discriminate between both scripts. 'Die Hard in Space' triggers strong commercial greenlights (99.9% Hit, 99.8% Franchise) while silencing the prestige channel (2.9%). Conversely, 'The Notebook 2' triggers a decisive prestige greenlight (99.9% Oscar) while failing the commercial thresholds.
The Dense Layer Network Graph
In standard neural network architectural diagrams, this algebraic pipeline maps directly to a directed computational graph:
Each input node represents an entry in the incoming feature vector x (x1,x2,x3,x4).
Each directed connection line between input node j and neuron i carries an individual weight parameter Wij from the weight matrix W.
Each neuron circle represents a linear sum accumulator computing zi=wi⋅x+bi.
Each output node emits the activated coordinate ai=σ(zi).
To visualize how incoming features propagate through the complete dense layer forward pass, trace the end-to-end computational pipeline below:
The complete dense layer forward pass pipeline ($x \to Wx \to +b \to z \to \sigma(z) \to a$): The 4-dimensional input vector $x$ is multiplied by the $3 \times 4$ weight matrix $W$ to produce 3 parallel dot products, shifted by the 3-channel bias vector $b$, and squashed coordinate-wise through activation functions to generate the final 3-neuron output vector.
Notice how the visual graph directly reflects the matrix dimensions. The 4 input nodes and 3 neuron circles require 3×4=12 connection lines—the exact count of entries in the weight matrix W. The 3 baseline biases in b inject directly into each neuron's summation node, and the squashing function σ operates independently at each exit port. This direct correspondence between visual nodes, connection lines, and matrix entries ensures that every architectural diagram in deep learning can be read as a concrete matrix-vector equation.
The Bridge to Deep Networks (Topic 4)
In this topic, our 3-neuron layer served as the final output of our network, producing predictions directly for the studio committee (a∈R3).
However, the true power of neural networks emerges when multiple dense layers are connected in sequence. The activation vector a(1)∈R3 produced by Layer 1 becomes the input vector directly entering Layer 2:
To Layer 2, a(1) is simply an ordered list of numbers. Layer 2 does not know or care how many raw features originally existed in x; it simply evaluates the features extracted by Layer 1.
In Topic 4, we will chain multiple weight matrices together to build Multi-Layer Perceptrons (MLPs), introducing hidden layers and discovering how non-linear activations like ReLU allow deep networks to solve problems a single layer can never crack.