The Complete MLP Forward Pass Math hero
Lesson 5Multi-Layer Neural Networks

The Complete MLP Forward Pass Math

Formulate the end-to-end forward pass equation chaining matrix transformations, biases, and activations from input vector to final predictions.

We are now ready to trace the complete, end-to-end mathematical forward pass through our Multi-Layer Perceptron.


The Composite Forward Pass Equation

The full forward pass of a 2-layer neural network is expressed as a single composite mathematical function:

a(2)=σ(W(2)f(W(1)x+b(1))+b(2))a^{(2)} = \sigma\left(W^{(2)} f\left(W^{(1)} x + b^{(1)}\right) + b^{(2)}\right)

Where:

  • x∈R4x \in \mathbb{R}^4 is the input feature vector.
  • W(1)∈R2×4W^{(1)} \in \mathbb{R}^{2 \times 4} and b(1)∈R2×1b^{(1)} \in \mathbb{R}^{2 \times 1} are Layer 1 parameters.
  • f(z)=ReLU(z)=max⁡(0,z)f(z) = \text{ReLU}(z) = \max(0, z) is the hidden layer activation function.
  • W(2)∈R1×2W^{(2)} \in \mathbb{R}^{1 \times 2} and b(2)∈R1×1b^{(2)} \in \mathbb{R}^{1 \times 1} are Layer 2 parameters.
  • σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}} is the output layer Sigmoid activation function.
  • a(2)=y^∈(0,1)a^{(2)} = \hat{y} \in (0, 1) is the final scalar predicted probability.

To execute this equation, we break it into four sequential computational steps:

Step 1 (Layer 1 Affine Transformation):z(1)=W(1)x+b(1)\mathbf{Step\ 1\ (Layer\ 1\ Affine\ Transformation):} \qquad z^{(1)} = W^{(1)} x + b^{(1)}

Step 2 (Layer 1 Non-Linear Activation):a(1)=ReLU(z(1))=max⁡(0,z(1))\mathbf{Step\ 2\ (Layer\ 1\ Non\text{-}Linear\ Activation):} \qquad a^{(1)} = \text{ReLU}\left(z^{(1)}\right) = \max\left(0, z^{(1)}\right)

Step 3 (Layer 2 Affine Transformation):z(2)=W(2)a(1)+b(2)\mathbf{Step\ 3\ (Layer\ 2\ Affine\ Transformation):} \qquad z^{(2)} = W^{(2)} a^{(1)} + b^{(2)}

Step 4 (Layer 2 Sigmoid Squashing):a(2)=σ(z(2))=11+e−z(2)\mathbf{Step\ 4\ (Layer\ 2\ Sigmoid\ Squashing):} \qquad a^{(2)} = \sigma\left(z^{(2)}\right) = \frac{1}{1 + e^{-z^{(2)}}}

To see how data flows sequentially through each computational card in this pipeline, examine the end-to-end dataflow diagram below:

Complete 2-Layer MLP Forward Pass Pipeline Flow
Complete 2-layer forward pass computational dataflow: Input vector x undergoes Layer 1 affine transformation ($W^{(1)}x + b^{(1)}$) to produce pre-activation $z^{(1)}$, squashed by ReLU into hidden representation vector $a^{(1)}$. Layer 2 computes affine sum $W^{(2)}a^{(1)} + b^{(2)}$ to produce scalar $z^{(2)}$, compressed through Sigmoid into final probability $a^{(2)}$.

Numerical Trace 1: 'Die Hard in Space'

Let's compute the exact forward pass for 'Die Hard in Space':

xDie Hard=[1.00.00.51.0](Action)(Romance)(Comedy)(Sci-Fi)x_{\text{Die Hard}} = \begin{bmatrix} 1.0 \\ 0.0 \\ 0.5 \\ 1.0 \end{bmatrix} \begin{matrix} \text{(Action)} \\ \text{(Romance)} \\ \text{(Comedy)} \\ \text{(Sci-Fi)} \end{matrix}

Step 1 & 2: Hidden Layer 1 Forward Calculation

Feature DimensionFeature Value (xx)Popcorn Neuron (W1j(1)W_{1j}^{(1)})Rom-Com Neuron (W2j(1)W_{2j}^{(1)})
Action (j=1j=1)1.01.01.0×3.0=+3.01.0 \times 3.0 = +3.01.0×−2.0=−2.01.0 \times -2.0 = -2.0
Romance (j=2j=2)0.00.00.0×−2.0=0.00.0 \times -2.0 = 0.00.0×3.0=0.00.0 \times 3.0 = 0.0
Comedy (j=3j=3)0.50.50.5×0.0=0.00.5 \times 0.0 = 0.00.5×2.0=+1.00.5 \times 2.0 = +1.0
Sci-Fi (j=4j=4)1.01.01.0×3.0=+3.01.0 \times 3.0 = +3.01.0×−1.0=−1.01.0 \times -1.0 = -1.0
Raw Dot Product (W(1)xW^{(1)} x)+6.0+6.0−2.0-2.0
Hidden Layer Bias (b(1)b^{(1)})−2.0-2.0−1.5-1.5
Pre-activation (z(1)=W(1)x+b(1)z^{(1)} = W^{(1)}x + b^{(1)})+4.0+4.0−3.5-3.5
ReLU Activation (a(1)=ReLU(z(1))a^{(1)} = \text{ReLU}(z^{(1)}))max⁡(0,4.0)=4.0\max(0, 4.0) = \mathbf{4.0}max⁡(0,−3.5)=0.0\max(0, -3.5) = \mathbf{0.0}

The intermediate hidden activation vector is:

aDie Hard(1)=[4.00.0](Popcorn Flick Factor)(Rom-Com Factor)a^{(1)}_{\text{Die Hard}} = \begin{bmatrix} 4.0 \\ 0.0 \end{bmatrix} \begin{matrix} \text{(Popcorn Flick Factor)} \\ \text{(Rom-Com Factor)} \end{matrix}

Notice how ReLU provides hard sparsity: because 'Die Hard in Space' produced a negative pre-activation on the Rom-Com neuron (z2(1)=−3.5z_2^{(1)} = -3.5), the neuron is completely silenced (a2(1)=0.0a_2^{(1)} = 0.0). The Popcorn Flick neuron activates with a strong positive signal (a1(1)=4.0a_1^{(1)} = 4.0).

Step 3 & 4: Output Layer 2 Forward Calculation

Now we propagate a(1)a^{(1)} through Layer 2:

z(2)=W(2)a(1)+b(2)=[1.51.0][4.00.0]+(−2.0)z^{(2)} = W^{(2)} a^{(1)} + b^{(2)} = \begin{bmatrix} 1.5 & 1.0 \end{bmatrix} \begin{bmatrix} 4.0 \\ 0.0 \end{bmatrix} + (-2.0) z(2)=(1.5×4.0)+(1.0×0.0)+(−2.0)=6.0+0.0−2.0=+4.0z^{(2)} = (1.5 \times 4.0) + (1.0 \times 0.0) + (-2.0) = 6.0 + 0.0 - 2.0 = \mathbf{+4.0}

Pass z(2)=+4.0z^{(2)} = +4.0 through the Sigmoid activation function:

a(2)=σ(4.0)=11+e−4.0=11+0.0183=11.0183≈0.9820  ⟹  98.2%a^{(2)} = \sigma(4.0) = \frac{1}{1 + e^{-4.0}} = \frac{1}{1 + 0.0183} = \frac{1}{1.0183} \approx \mathbf{0.9820} \implies \mathbf{98.2\%}

'Die Hard in Space' receives an overwhelming 98.2%98.2\% Hit Probability.


Numerical Trace 2: 'The Notebook 2'

Now let's trace our second script, the romantic sequel 'The Notebook 2':

xNotebook 2=[0.01.00.50.0](Action)(Romance)(Comedy)(Sci-Fi)x_{\text{Notebook 2}} = \begin{bmatrix} 0.0 \\ 1.0 \\ 0.5 \\ 0.0 \end{bmatrix} \begin{matrix} \text{(Action)} \\ \text{(Romance)} \\ \text{(Comedy)} \\ \text{(Sci-Fi)} \end{matrix}

Step 1 & 2: Hidden Layer 1 Forward Calculation

Feature DimensionFeature Value (xx)Popcorn Neuron (W1j(1)W_{1j}^{(1)})Rom-Com Neuron (W2j(1)W_{2j}^{(1)})
Action (j=1j=1)0.00.00.0×3.0=0.00.0 \times 3.0 = 0.00.0×−2.0=0.00.0 \times -2.0 = 0.0
Romance (j=2j=2)1.01.01.0×−2.0=−2.01.0 \times -2.0 = -2.01.0×3.0=+3.01.0 \times 3.0 = +3.0
Comedy (j=3j=3)0.50.50.5×0.0=0.00.5 \times 0.0 = 0.00.5×2.0=+1.00.5 \times 2.0 = +1.0
Sci-Fi (j=4j=4)0.00.00.0×3.0=0.00.0 \times 3.0 = 0.00.0×−1.0=0.00.0 \times -1.0 = 0.0
Raw Dot Product (W(1)xW^{(1)} x)−2.0-2.0+4.0+4.0
Hidden Layer Bias (b(1)b^{(1)})−2.0-2.0−1.5-1.5
Pre-activation (z(1)=W(1)x+b(1)z^{(1)} = W^{(1)}x + b^{(1)})−4.0-4.0+2.5+2.5
ReLU Activation (a(1)=ReLU(z(1))a^{(1)} = \text{ReLU}(z^{(1)}))max⁡(0,−4.0)=0.0\max(0, -4.0) = \mathbf{0.0}max⁡(0,2.5)=2.5\max(0, 2.5) = \mathbf{2.5}

The intermediate hidden activation vector is:

aNotebook 2(1)=[0.02.5](Popcorn Flick Factor)(Rom-Com Factor)a^{(1)}_{\text{Notebook 2}} = \begin{bmatrix} 0.0 \\ 2.5 \end{bmatrix} \begin{matrix} \text{(Popcorn Flick Factor)} \\ \text{(Rom-Com Factor)} \end{matrix}

Here, the roles are reversed: the Popcorn Flick neuron is completely silenced (0.00.0), while the Rom-Com neuron fires strongly (2.52.5).

Step 3 & 4: Output Layer 2 Forward Calculation

Propagate a(1)a^{(1)} through Layer 2:

z(2)=W(2)a(1)+b(2)=[1.51.0][0.02.5]+(−2.0)z^{(2)} = W^{(2)} a^{(1)} + b^{(2)} = \begin{bmatrix} 1.5 & 1.0 \end{bmatrix} \begin{bmatrix} 0.0 \\ 2.5 \end{bmatrix} + (-2.0) z(2)=(1.5×0.0)+(1.0×2.5)+(−2.0)=0.0+2.5−2.0=+0.5z^{(2)} = (1.5 \times 0.0) + (1.0 \times 2.5) + (-2.0) = 0.0 + 2.5 - 2.0 = \mathbf{+0.5}

Pass z(2)=+0.5z^{(2)} = +0.5 through the Sigmoid activation function:

a(2)=σ(0.5)=11+e−0.5=11+0.6065=11.6065≈0.6225  ⟹  62.3%a^{(2)} = \sigma(0.5) = \frac{1}{1 + e^{-0.5}} = \frac{1}{1 + 0.6065} = \frac{1}{1.6065} \approx \mathbf{0.6225} \implies \mathbf{62.3\%}

'The Notebook 2' overcomes the studio market hurdle to earn a solid 62.3%62.3\% Greenlight Probability.


The Complete Forward Pass Milestone

We have now assembled the entire forward execution engine of deep learning:

[ Raw Numbers ] ──► [ Dot Products ] ──► [ Biases ] ──► [ Activations ] ──► [ Matrices ] ──► [ Hidden Layers ] ──► [ Prediction ]
  1. Topic 1: We encoded qualitative reality into 1D Feature Vectors (xx) and computed raw alignment using Dot Products (w⋅xw \cdot x).
  2. Topic 2: We added baseline Biases (+b+b) to create affine sums (z=w⋅x+bz = w \cdot x + b) and applied non-linear Activation Functions (σ,ReLU\sigma, \text{ReLU}) to establish decision boundaries.
  3. Topic 3: We stacked weight vectors into Matrices (WW) to evaluate multiple parallel decision channels (Layer Width).
  4. Topic 4: We stacked layers into an MLP (Depth), using Hidden Layers and Representation Learning to construct abstract concept spaces that make non-linear classification possible.

Data enters as raw feature values, flows through geometric transformations, and emerges as a calibrated, probabilistic prediction.


Previous
The Rectified Linear Unit Activation