We are now ready to trace the complete, end-to-end mathematical forward pass through our Multi-Layer Perceptron.
The Composite Forward Pass Equation
The full forward pass of a 2-layer neural network is expressed as a single composite mathematical function:
a ( 2 ) = σ ( W ( 2 ) f ( W ( 1 ) x + b ( 1 ) ) + b ( 2 ) ) a^{(2)} = \sigma\left(W^{(2)} f\left(W^{(1)} x + b^{(1)}\right) + b^{(2)}\right) a ( 2 ) = σ ( W ( 2 ) f ( W ( 1 ) x + b ( 1 ) ) + b ( 2 ) )
Where:
x ∈ R 4 x \in \mathbb{R}^4 x ∈ R 4 is the input feature vector.
W ( 1 ) ∈ R 2 × 4 W^{(1)} \in \mathbb{R}^{2 \times 4} W ( 1 ) ∈ R 2 × 4 and b ( 1 ) ∈ R 2 × 1 b^{(1)} \in \mathbb{R}^{2 \times 1} b ( 1 ) ∈ R 2 × 1 are Layer 1 parameters.
f ( z ) = ReLU ( z ) = max ( 0 , z ) f(z) = \text{ReLU}(z) = \max(0, z) f ( z ) = ReLU ( z ) = max ( 0 , z ) is the hidden layer activation function.
W ( 2 ) ∈ R 1 × 2 W^{(2)} \in \mathbb{R}^{1 \times 2} W ( 2 ) ∈ R 1 × 2 and b ( 2 ) ∈ R 1 × 1 b^{(2)} \in \mathbb{R}^{1 \times 1} b ( 2 ) ∈ R 1 × 1 are Layer 2 parameters.
σ ( z ) = 1 1 + e − z \sigma(z) = \frac{1}{1 + e^{-z}} σ ( z ) = 1 + e − z 1 is the output layer Sigmoid activation function.
a ( 2 ) = y ^ ∈ ( 0 , 1 ) a^{(2)} = \hat{y} \in (0, 1) a ( 2 ) = y ^ ∈ ( 0 , 1 ) is the final scalar predicted probability.
To execute this equation, we break it into four sequential computational steps:
S t e p 1 ( L a y e r 1 A f f i n e T r a n s f o r m a t i o n ) : z ( 1 ) = W ( 1 ) x + b ( 1 ) \mathbf{Step\ 1\ (Layer\ 1\ Affine\ Transformation):} \qquad z^{(1)} = W^{(1)} x + b^{(1)} Step 1 ( Layer 1 Affine Transformation ) : z ( 1 ) = W ( 1 ) x + b ( 1 )
S t e p 2 ( L a y e r 1 N o n - L i n e a r A c t i v a t i o n ) : a ( 1 ) = ReLU ( z ( 1 ) ) = max ( 0 , z ( 1 ) ) \mathbf{Step\ 2\ (Layer\ 1\ Non\text{-}Linear\ Activation):} \qquad a^{(1)} = \text{ReLU}\left(z^{(1)}\right) = \max\left(0, z^{(1)}\right) Step 2 ( Layer 1 Non - Linear Activation ) : a ( 1 ) = ReLU ( z ( 1 ) ) = max ( 0 , z ( 1 ) )
S t e p 3 ( L a y e r 2 A f f i n e T r a n s f o r m a t i o n ) : z ( 2 ) = W ( 2 ) a ( 1 ) + b ( 2 ) \mathbf{Step\ 3\ (Layer\ 2\ Affine\ Transformation):} \qquad z^{(2)} = W^{(2)} a^{(1)} + b^{(2)} Step 3 ( Layer 2 Affine Transformation ) : z ( 2 ) = W ( 2 ) a ( 1 ) + b ( 2 )
S t e p 4 ( L a y e r 2 S i g m o i d S q u a s h i n g ) : a ( 2 ) = σ ( z ( 2 ) ) = 1 1 + e − z ( 2 ) \mathbf{Step\ 4\ (Layer\ 2\ Sigmoid\ Squashing):} \qquad a^{(2)} = \sigma\left(z^{(2)}\right) = \frac{1}{1 + e^{-z^{(2)}}} Step 4 ( Layer 2 Sigmoid Squashing ) : a ( 2 ) = σ ( z ( 2 ) ) = 1 + e − z ( 2 ) 1
To see how data flows sequentially through each computational card in this pipeline, examine the end-to-end dataflow diagram below:
Complete 2-layer forward pass computational dataflow: Input vector x undergoes Layer 1 affine transformation ($W^{(1)}x + b^{(1)}$) to produce pre-activation $z^{(1)}$, squashed by ReLU into hidden representation vector $a^{(1)}$. Layer 2 computes affine sum $W^{(2)}a^{(1)} + b^{(2)}$ to produce scalar $z^{(2)}$, compressed through Sigmoid into final probability $a^{(2)}$.
Numerical Trace 1: 'Die Hard in Space'
Let's compute the exact forward pass for 'Die Hard in Space' :
x Die Hard = [ 1.0 0.0 0.5 1.0 ] (Action) (Romance) (Comedy) (Sci-Fi) x_{\text{Die Hard}} = \begin{bmatrix} 1.0 \\ 0.0 \\ 0.5 \\ 1.0 \end{bmatrix} \begin{matrix} \text{(Action)} \\ \text{(Romance)} \\ \text{(Comedy)} \\ \text{(Sci-Fi)} \end{matrix} x Die Hard = 1.0 0.0 0.5 1.0 (Action) (Romance) (Comedy) (Sci-Fi)
Step 1 & 2: Hidden Layer 1 Forward Calculation
Feature Dimension Feature Value (x x x ) Popcorn Neuron (W 1 j ( 1 ) W_{1j}^{(1)} W 1 j ( 1 ) ) Rom-Com Neuron (W 2 j ( 1 ) W_{2j}^{(1)} W 2 j ( 1 ) ) Action (j = 1 j=1 j = 1 )1.0 1.0 1.0 1.0 × 3.0 = + 3.0 1.0 \times 3.0 = +3.0 1.0 × 3.0 = + 3.0 1.0 × − 2.0 = − 2.0 1.0 \times -2.0 = -2.0 1.0 × − 2.0 = − 2.0 Romance (j = 2 j=2 j = 2 )0.0 0.0 0.0 0.0 × − 2.0 = 0.0 0.0 \times -2.0 = 0.0 0.0 × − 2.0 = 0.0 0.0 × 3.0 = 0.0 0.0 \times 3.0 = 0.0 0.0 × 3.0 = 0.0 Comedy (j = 3 j=3 j = 3 )0.5 0.5 0.5 0.5 × 0.0 = 0.0 0.5 \times 0.0 = 0.0 0.5 × 0.0 = 0.0 0.5 × 2.0 = + 1.0 0.5 \times 2.0 = +1.0 0.5 × 2.0 = + 1.0 Sci-Fi (j = 4 j=4 j = 4 )1.0 1.0 1.0 1.0 × 3.0 = + 3.0 1.0 \times 3.0 = +3.0 1.0 × 3.0 = + 3.0 1.0 × − 1.0 = − 1.0 1.0 \times -1.0 = -1.0 1.0 × − 1.0 = − 1.0 Raw Dot Product (W ( 1 ) x W^{(1)} x W ( 1 ) x ) + 6.0 +6.0 + 6.0 − 2.0 -2.0 − 2.0 Hidden Layer Bias (b ( 1 ) b^{(1)} b ( 1 ) ) − 2.0 -2.0 − 2.0 − 1.5 -1.5 − 1.5 Pre-activation (z ( 1 ) = W ( 1 ) x + b ( 1 ) z^{(1)} = W^{(1)}x + b^{(1)} z ( 1 ) = W ( 1 ) x + b ( 1 ) ) + 4.0 +4.0 + 4.0 − 3.5 -3.5 − 3.5 ReLU Activation (a ( 1 ) = ReLU ( z ( 1 ) ) a^{(1)} = \text{ReLU}(z^{(1)}) a ( 1 ) = ReLU ( z ( 1 ) ) ) max ( 0 , 4.0 ) = 4.0 \max(0, 4.0) = \mathbf{4.0} max ( 0 , 4.0 ) = 4.0 max ( 0 , − 3.5 ) = 0.0 \max(0, -3.5) = \mathbf{0.0} max ( 0 , − 3.5 ) = 0.0
The intermediate hidden activation vector is:
a Die Hard ( 1 ) = [ 4.0 0.0 ] (Popcorn Flick Factor) (Rom-Com Factor) a^{(1)}_{\text{Die Hard}} = \begin{bmatrix} 4.0 \\ 0.0 \end{bmatrix} \begin{matrix} \text{(Popcorn Flick Factor)} \\ \text{(Rom-Com Factor)} \end{matrix} a Die Hard ( 1 ) = [ 4.0 0.0 ] (Popcorn Flick Factor) (Rom-Com Factor)
Notice how ReLU provides hard sparsity : because 'Die Hard in Space' produced a negative pre-activation on the Rom-Com neuron (z 2 ( 1 ) = − 3.5 z_2^{(1)} = -3.5 z 2 ( 1 ) = − 3.5 ), the neuron is completely silenced (a 2 ( 1 ) = 0.0 a_2^{(1)} = 0.0 a 2 ( 1 ) = 0.0 ). The Popcorn Flick neuron activates with a strong positive signal (a 1 ( 1 ) = 4.0 a_1^{(1)} = 4.0 a 1 ( 1 ) = 4.0 ).
Step 3 & 4: Output Layer 2 Forward Calculation
Now we propagate a ( 1 ) a^{(1)} a ( 1 ) through Layer 2:
z ( 2 ) = W ( 2 ) a ( 1 ) + b ( 2 ) = [ 1.5 1.0 ] [ 4.0 0.0 ] + ( − 2.0 ) z^{(2)} = W^{(2)} a^{(1)} + b^{(2)} = \begin{bmatrix} 1.5 & 1.0 \end{bmatrix} \begin{bmatrix} 4.0 \\ 0.0 \end{bmatrix} + (-2.0) z ( 2 ) = W ( 2 ) a ( 1 ) + b ( 2 ) = [ 1.5 1.0 ] [ 4.0 0.0 ] + ( − 2.0 )
z ( 2 ) = ( 1.5 × 4.0 ) + ( 1.0 × 0.0 ) + ( − 2.0 ) = 6.0 + 0.0 − 2.0 = + 4.0 z^{(2)} = (1.5 \times 4.0) + (1.0 \times 0.0) + (-2.0) = 6.0 + 0.0 - 2.0 = \mathbf{+4.0} z ( 2 ) = ( 1.5 × 4.0 ) + ( 1.0 × 0.0 ) + ( − 2.0 ) = 6.0 + 0.0 − 2.0 = + 4.0
Pass z ( 2 ) = + 4.0 z^{(2)} = +4.0 z ( 2 ) = + 4.0 through the Sigmoid activation function:
a ( 2 ) = σ ( 4.0 ) = 1 1 + e − 4.0 = 1 1 + 0.0183 = 1 1.0183 ≈ 0.9820 ⟹ 98.2 % a^{(2)} = \sigma(4.0) = \frac{1}{1 + e^{-4.0}} = \frac{1}{1 + 0.0183} = \frac{1}{1.0183} \approx \mathbf{0.9820} \implies \mathbf{98.2\%} a ( 2 ) = σ ( 4.0 ) = 1 + e − 4.0 1 = 1 + 0.0183 1 = 1.0183 1 ≈ 0.9820 ⟹ 98.2%
'Die Hard in Space' receives an overwhelming 98.2 % 98.2\% 98.2% Hit Probability .
Numerical Trace 2: 'The Notebook 2'
Now let's trace our second script, the romantic sequel 'The Notebook 2' :
x Notebook 2 = [ 0.0 1.0 0.5 0.0 ] (Action) (Romance) (Comedy) (Sci-Fi) x_{\text{Notebook 2}} = \begin{bmatrix} 0.0 \\ 1.0 \\ 0.5 \\ 0.0 \end{bmatrix} \begin{matrix} \text{(Action)} \\ \text{(Romance)} \\ \text{(Comedy)} \\ \text{(Sci-Fi)} \end{matrix} x Notebook 2 = 0.0 1.0 0.5 0.0 (Action) (Romance) (Comedy) (Sci-Fi)
Step 1 & 2: Hidden Layer 1 Forward Calculation
Feature Dimension Feature Value (x x x ) Popcorn Neuron (W 1 j ( 1 ) W_{1j}^{(1)} W 1 j ( 1 ) ) Rom-Com Neuron (W 2 j ( 1 ) W_{2j}^{(1)} W 2 j ( 1 ) ) Action (j = 1 j=1 j = 1 )0.0 0.0 0.0 0.0 × 3.0 = 0.0 0.0 \times 3.0 = 0.0 0.0 × 3.0 = 0.0 0.0 × − 2.0 = 0.0 0.0 \times -2.0 = 0.0 0.0 × − 2.0 = 0.0 Romance (j = 2 j=2 j = 2 )1.0 1.0 1.0 1.0 × − 2.0 = − 2.0 1.0 \times -2.0 = -2.0 1.0 × − 2.0 = − 2.0 1.0 × 3.0 = + 3.0 1.0 \times 3.0 = +3.0 1.0 × 3.0 = + 3.0 Comedy (j = 3 j=3 j = 3 )0.5 0.5 0.5 0.5 × 0.0 = 0.0 0.5 \times 0.0 = 0.0 0.5 × 0.0 = 0.0 0.5 × 2.0 = + 1.0 0.5 \times 2.0 = +1.0 0.5 × 2.0 = + 1.0 Sci-Fi (j = 4 j=4 j = 4 )0.0 0.0 0.0 0.0 × 3.0 = 0.0 0.0 \times 3.0 = 0.0 0.0 × 3.0 = 0.0 0.0 × − 1.0 = 0.0 0.0 \times -1.0 = 0.0 0.0 × − 1.0 = 0.0 Raw Dot Product (W ( 1 ) x W^{(1)} x W ( 1 ) x ) − 2.0 -2.0 − 2.0 + 4.0 +4.0 + 4.0 Hidden Layer Bias (b ( 1 ) b^{(1)} b ( 1 ) ) − 2.0 -2.0 − 2.0 − 1.5 -1.5 − 1.5 Pre-activation (z ( 1 ) = W ( 1 ) x + b ( 1 ) z^{(1)} = W^{(1)}x + b^{(1)} z ( 1 ) = W ( 1 ) x + b ( 1 ) ) − 4.0 -4.0 − 4.0 + 2.5 +2.5 + 2.5 ReLU Activation (a ( 1 ) = ReLU ( z ( 1 ) ) a^{(1)} = \text{ReLU}(z^{(1)}) a ( 1 ) = ReLU ( z ( 1 ) ) ) max ( 0 , − 4.0 ) = 0.0 \max(0, -4.0) = \mathbf{0.0} max ( 0 , − 4.0 ) = 0.0 max ( 0 , 2.5 ) = 2.5 \max(0, 2.5) = \mathbf{2.5} max ( 0 , 2.5 ) = 2.5
The intermediate hidden activation vector is:
a Notebook 2 ( 1 ) = [ 0.0 2.5 ] (Popcorn Flick Factor) (Rom-Com Factor) a^{(1)}_{\text{Notebook 2}} = \begin{bmatrix} 0.0 \\ 2.5 \end{bmatrix} \begin{matrix} \text{(Popcorn Flick Factor)} \\ \text{(Rom-Com Factor)} \end{matrix} a Notebook 2 ( 1 ) = [ 0.0 2.5 ] (Popcorn Flick Factor) (Rom-Com Factor)
Here, the roles are reversed: the Popcorn Flick neuron is completely silenced (0.0 0.0 0.0 ), while the Rom-Com neuron fires strongly (2.5 2.5 2.5 ).
Step 3 & 4: Output Layer 2 Forward Calculation
Propagate a ( 1 ) a^{(1)} a ( 1 ) through Layer 2:
z ( 2 ) = W ( 2 ) a ( 1 ) + b ( 2 ) = [ 1.5 1.0 ] [ 0.0 2.5 ] + ( − 2.0 ) z^{(2)} = W^{(2)} a^{(1)} + b^{(2)} = \begin{bmatrix} 1.5 & 1.0 \end{bmatrix} \begin{bmatrix} 0.0 \\ 2.5 \end{bmatrix} + (-2.0) z ( 2 ) = W ( 2 ) a ( 1 ) + b ( 2 ) = [ 1.5 1.0 ] [ 0.0 2.5 ] + ( − 2.0 )
z ( 2 ) = ( 1.5 × 0.0 ) + ( 1.0 × 2.5 ) + ( − 2.0 ) = 0.0 + 2.5 − 2.0 = + 0.5 z^{(2)} = (1.5 \times 0.0) + (1.0 \times 2.5) + (-2.0) = 0.0 + 2.5 - 2.0 = \mathbf{+0.5} z ( 2 ) = ( 1.5 × 0.0 ) + ( 1.0 × 2.5 ) + ( − 2.0 ) = 0.0 + 2.5 − 2.0 = + 0.5
Pass z ( 2 ) = + 0.5 z^{(2)} = +0.5 z ( 2 ) = + 0.5 through the Sigmoid activation function:
a ( 2 ) = σ ( 0.5 ) = 1 1 + e − 0.5 = 1 1 + 0.6065 = 1 1.6065 ≈ 0.6225 ⟹ 62.3 % a^{(2)} = \sigma(0.5) = \frac{1}{1 + e^{-0.5}} = \frac{1}{1 + 0.6065} = \frac{1}{1.6065} \approx \mathbf{0.6225} \implies \mathbf{62.3\%} a ( 2 ) = σ ( 0.5 ) = 1 + e − 0.5 1 = 1 + 0.6065 1 = 1.6065 1 ≈ 0.6225 ⟹ 62.3%
'The Notebook 2' overcomes the studio market hurdle to earn a solid 62.3 % 62.3\% 62.3% Greenlight Probability .
The Complete Forward Pass Milestone
We have now assembled the entire forward execution engine of deep learning:
[ Raw Numbers ] ──► [ Dot Products ] ──► [ Biases ] ──► [ Activations ] ──► [ Matrices ] ──► [ Hidden Layers ] ──► [ Prediction ]
Topic 1: We encoded qualitative reality into 1D Feature Vectors (x x x ) and computed raw alignment using Dot Products (w ⋅ x w \cdot x w ⋅ x ).
Topic 2: We added baseline Biases (+ b +b + b ) to create affine sums (z = w ⋅ x + b z = w \cdot x + b z = w ⋅ x + b ) and applied non-linear Activation Functions (σ , ReLU \sigma, \text{ReLU} σ , ReLU ) to establish decision boundaries.
Topic 3: We stacked weight vectors into Matrices (W W W ) to evaluate multiple parallel decision channels (Layer Width ).
Topic 4: We stacked layers into an MLP (Depth ), using Hidden Layers and Representation Learning to construct abstract concept spaces that make non-linear classification possible.
Data enters as raw feature values, flows through geometric transformations, and emerges as a calibrated, probabilistic prediction.