Work through every problem on paper before expanding the step-by-step solutions.
Part 1: Abstract Mechanics
Work through these 5 rapid-fire drill problems testing tensor flow, forward passes, loss evaluation, output attribution, and hidden backpropagation.
Problem 1: End-to-End Dimension Flow Audit
Consider a 3-layer Multi-Layer Perceptron (d 0 → d 1 → d 2 → d 3 d_0 \to d_1 \to d_2 \to d_3 d 0 → d 1 → d 2 → d 3 ) with layer widths d 0 = 3 d_0 = 3 d 0 = 3 (inputs), d 1 = 4 d_1 = 4 d 1 = 4 (hidden layer 1), d 2 = 2 d_2 = 2 d 2 = 2 (hidden layer 2), and d 3 = 1 d_3 = 1 d 3 = 1 (output scalar).
State the exact mathematical matrix dimensions for each of the following tensors:
Input vector x x x and transposed input x T x^T x T .
Weight matrices W ( 1 ) , W ( 2 ) , W ( 3 ) W^{(1)}, W^{(2)}, W^{(3)} W ( 1 ) , W ( 2 ) , W ( 3 ) and bias vectors b ( 1 ) , b ( 2 ) , b ( 3 ) b^{(1)}, b^{(2)}, b^{(3)} b ( 1 ) , b ( 2 ) , b ( 3 ) .
Pre-activation vectors z ( 1 ) , z ( 2 ) , z ( 3 ) z^{(1)}, z^{(2)}, z^{(3)} z ( 1 ) , z ( 2 ) , z ( 3 ) and activation vectors a ( 1 ) , a ( 2 ) , a ( 3 ) a^{(1)}, a^{(2)}, a^{(3)} a ( 1 ) , a ( 2 ) , a ( 3 ) .
Error deltas δ ( 3 ) , δ ( 2 ) , δ ( 1 ) \delta^{(3)}, \delta^{(2)}, \delta^{(1)} δ ( 3 ) , δ ( 2 ) , δ ( 1 ) .
Parameter gradient matrices ∂ L ∂ W ( 3 ) , ∂ L ∂ W ( 2 ) , ∂ L ∂ W ( 1 ) \frac{\partial L}{\partial W^{(3)}}, \frac{\partial L}{\partial W^{(2)}}, \frac{\partial L}{\partial W^{(1)}} ∂ W ( 3 ) ∂ L , ∂ W ( 2 ) ∂ L , ∂ W ( 1 ) ∂ L .
Reveal Solution
Input Tensors:
x ∈ R 3 × 1 x \in \mathbb{R}^{3 \times 1} x ∈ R 3 × 1
x T ∈ R 1 × 3 x^T \in \mathbb{R}^{1 \times 3} x T ∈ R 1 × 3
Parameter Tensors (d out × d in d_{\text{out}} \times d_{\text{in}} d out × d in ):
W ( 1 ) ∈ R 4 × 3 , b ( 1 ) ∈ R 4 × 1 W^{(1)} \in \mathbb{R}^{4 \times 3}, \quad b^{(1)} \in \mathbb{R}^{4 \times 1} W ( 1 ) ∈ R 4 × 3 , b ( 1 ) ∈ R 4 × 1
W ( 2 ) ∈ R 2 × 4 , b ( 2 ) ∈ R 2 × 1 W^{(2)} \in \mathbb{R}^{2 \times 4}, \quad b^{(2)} \in \mathbb{R}^{2 \times 1} W ( 2 ) ∈ R 2 × 4 , b ( 2 ) ∈ R 2 × 1
W ( 3 ) ∈ R 1 × 2 , b ( 3 ) ∈ R 1 × 1 W^{(3)} \in \mathbb{R}^{1 \times 2}, \quad b^{(3)} \in \mathbb{R}^{1 \times 1} W ( 3 ) ∈ R 1 × 2 , b ( 3 ) ∈ R 1 × 1
Forward Activation Tensors:
z ( 1 ) , a ( 1 ) ∈ R 4 × 1 z^{(1)}, a^{(1)} \in \mathbb{R}^{4 \times 1} z ( 1 ) , a ( 1 ) ∈ R 4 × 1
z ( 2 ) , a ( 2 ) ∈ R 2 × 1 z^{(2)}, a^{(2)} \in \mathbb{R}^{2 \times 1} z ( 2 ) , a ( 2 ) ∈ R 2 × 1
z ( 3 ) , a ( 3 ) ∈ R 1 × 1 z^{(3)}, a^{(3)} \in \mathbb{R}^{1 \times 1} z ( 3 ) , a ( 3 ) ∈ R 1 × 1 (scalar prediction y ^ \hat{y} y ^ )
Backward Error Deltas (δ ( l ) = ∂ L ∂ z ( l ) \delta^{(l)} = \frac{\partial L}{\partial z^{(l)}} δ ( l ) = ∂ z ( l ) ∂ L ):
δ ( 3 ) ∈ R 1 × 1 \delta^{(3)} \in \mathbb{R}^{1 \times 1} δ ( 3 ) ∈ R 1 × 1
δ ( 2 ) ∈ R 2 × 1 \delta^{(2)} \in \mathbb{R}^{2 \times 1} δ ( 2 ) ∈ R 2 × 1
δ ( 1 ) ∈ R 4 × 1 \delta^{(1)} \in \mathbb{R}^{4 \times 1} δ ( 1 ) ∈ R 4 × 1
Parameter Gradients (∂ L ∂ W ( l ) = δ ( l ) ( a ( l − 1 ) ) T \frac{\partial L}{\partial W^{(l)}} = \delta^{(l)} (a^{(l-1)})^T ∂ W ( l ) ∂ L = δ ( l ) ( a ( l − 1 ) ) T ):
∂ L ∂ W ( 3 ) = δ ( 3 ) ( a ( 2 ) ) T ∈ R 1 × 1 × R 1 × 2 = R 1 × 2 \frac{\partial L}{\partial W^{(3)}} = \delta^{(3)} (a^{(2)})^T \in \mathbb{R}^{1 \times 1} \times \mathbb{R}^{1 \times 2} = \mathbf{\mathbb{R}^{1 \times 2}} ∂ W ( 3 ) ∂ L = δ ( 3 ) ( a ( 2 ) ) T ∈ R 1 × 1 × R 1 × 2 = R 1 × 2
∂ L ∂ W ( 2 ) = δ ( 2 ) ( a ( 1 ) ) T ∈ R 2 × 1 × R 1 × 4 = R 2 × 4 \frac{\partial L}{\partial W^{(2)}} = \delta^{(2)} (a^{(1)})^T \in \mathbb{R}^{2 \times 1} \times \mathbb{R}^{1 \times 4} = \mathbf{\mathbb{R}^{2 \times 4}} ∂ W ( 2 ) ∂ L = δ ( 2 ) ( a ( 1 ) ) T ∈ R 2 × 1 × R 1 × 4 = R 2 × 4
∂ L ∂ W ( 1 ) = δ ( 1 ) x T ∈ R 4 × 1 × R 1 × 3 = R 4 × 3 \frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T \in \mathbb{R}^{4 \times 1} \times \mathbb{R}^{1 \times 3} = \mathbf{\mathbb{R}^{4 \times 3}} ∂ W ( 1 ) ∂ L = δ ( 1 ) x T ∈ R 4 × 1 × R 1 × 3 = R 4 × 3
Verification: Every parameter gradient matrix ∂ L ∂ W ( l ) \frac{\partial L}{\partial W^{(l)}} ∂ W ( l ) ∂ L matches the exact dimensions of its corresponding forward weight matrix W ( l ) W^{(l)} W ( l ) .
Problem 2: Complete 2 → 2 → 1 2 \to 2 \to 1 2 → 2 → 1 Forward Pass Calculation
A 2 → 2 → 1 2 \to 2 \to 1 2 → 2 → 1 Multi-Layer Perceptron has the following parameters:
Input vector: x = [ 2.0 1.0 ] x = \begin{bmatrix} 2.0 \\ 1.0 \end{bmatrix} x = [ 2.0 1.0 ]
Layer 1: W ( 1 ) = [ 1.0 2.0 − 1.0 1.0 ] , b ( 1 ) = [ − 1.0 0.5 ] W^{(1)} = \begin{bmatrix} 1.0 & 2.0 \\ -1.0 & 1.0 \end{bmatrix}, \quad b^{(1)} = \begin{bmatrix} -1.0 \\ 0.5 \end{bmatrix} W ( 1 ) = [ 1.0 − 1.0 2.0 1.0 ] , b ( 1 ) = [ − 1.0 0.5 ] , with ReLU activation.
Layer 2: W ( 2 ) = [ 2.0 1.0 ] , b ( 2 ) = [ − 1.0 ] W^{(2)} = \begin{bmatrix} 2.0 & 1.0 \end{bmatrix}, \quad b^{(2)} = \begin{bmatrix} -1.0 \end{bmatrix} W ( 2 ) = [ 2.0 1.0 ] , b ( 2 ) = [ − 1.0 ] , with Sigmoid activation.
Compute:
Layer 1 pre-activation vector z ( 1 ) z^{(1)} z ( 1 ) and hidden activation vector a ( 1 ) a^{(1)} a ( 1 ) .
Layer 2 pre-activation scalar z ( 2 ) z^{(2)} z ( 2 ) and output prediction y ^ = a ( 2 ) \hat{y} = a^{(2)} y ^ = a ( 2 ) (use e − 5.0 ≈ 0.006738 e^{-5.0} \approx 0.006738 e − 5.0 ≈ 0.006738 ).
Reveal Solution
Layer 1 Forward Calculation:
z 1 ( 1 ) = ( 1.0 ) ( 2.0 ) + ( 2.0 ) ( 1.0 ) − 1.0 = 2.0 + 2.0 − 1.0 = + 3.0 z_1^{(1)} = (1.0)(2.0) + (2.0)(1.0) - 1.0 = 2.0 + 2.0 - 1.0 = \mathbf{+3.0} z 1 ( 1 ) = ( 1.0 ) ( 2.0 ) + ( 2.0 ) ( 1.0 ) − 1.0 = 2.0 + 2.0 − 1.0 = + 3.0
z 2 ( 1 ) = ( − 1.0 ) ( 2.0 ) + ( 1.0 ) ( 1.0 ) + 0.5 = − 2.0 + 1.0 + 0.5 = − 0.5 z_2^{(1)} = (-1.0)(2.0) + (1.0)(1.0) + 0.5 = -2.0 + 1.0 + 0.5 = \mathbf{-0.5} z 2 ( 1 ) = ( − 1.0 ) ( 2.0 ) + ( 1.0 ) ( 1.0 ) + 0.5 = − 2.0 + 1.0 + 0.5 = − 0.5
z ( 1 ) = [ + 3.0 − 0.5 ] z^{(1)} = \begin{bmatrix} +3.0 \\ -0.5 \end{bmatrix} z ( 1 ) = [ + 3.0 − 0.5 ]
Apply ReLU:
a ( 1 ) = [ max ( 0 , 3.0 ) max ( 0 , − 0.5 ) ] = [ 3.0 0.0 ] a^{(1)} = \begin{bmatrix} \max(0, 3.0) \\ \max(0, -0.5) \end{bmatrix} = \begin{bmatrix} \mathbf{3.0} \\ \mathbf{0.0} \end{bmatrix} a ( 1 ) = [ max ( 0 , 3.0 ) max ( 0 , − 0.5 ) ] = [ 3.0 0.0 ]
Layer 2 Forward Calculation:
z ( 2 ) = W ( 2 ) a ( 1 ) + b ( 2 ) = ( 2.0 ) ( 3.0 ) + ( 1.0 ) ( 0.0 ) − 1.0 = 6.0 + 0.0 − 1.0 = + 5.0 z^{(2)} = W^{(2)} a^{(1)} + b^{(2)} = (2.0)(3.0) + (1.0)(0.0) - 1.0 = 6.0 + 0.0 - 1.0 = \mathbf{+5.0} z ( 2 ) = W ( 2 ) a ( 1 ) + b ( 2 ) = ( 2.0 ) ( 3.0 ) + ( 1.0 ) ( 0.0 ) − 1.0 = 6.0 + 0.0 − 1.0 = + 5.0
Apply Sigmoid:
a ( 2 ) = σ ( 5.0 ) = 1 1 + e − 5.0 = 1 1 + 0.006738 = 1 1.006738 ≈ 0.993307 ( 99.33 % ) a^{(2)} = \sigma(5.0) = \frac{1}{1 + e^{-5.0}} = \frac{1}{1 + 0.006738} = \frac{1}{1.006738} \approx \mathbf{0.993307} \quad (\mathbf{99.33\%}) a ( 2 ) = σ ( 5.0 ) = 1 + e − 5.0 1 = 1 + 0.006738 1 = 1.006738 1 ≈ 0.993307 ( 99.33% )
Problem 3: Mean Squared Error Loss & Sensitivity
A neural network outputs prediction y ^ = 0.80 \hat{y} = 0.80 y ^ = 0.80 for a training instance where ground truth is y = 0.0 y = 0.0 y = 0.0 .
Compute:
Raw prediction error e = y − y ^ e = y - \hat{y} e = y − y ^ .
Standard Mean Squared Error loss L = ( y − y ^ ) 2 L = (y - \hat{y})^2 L = ( y − y ^ ) 2 .
Scaled MSE loss L scaled = 1 2 ( y − y ^ ) 2 L_{\text{scaled}} = \frac{1}{2}(y - \hat{y})^2 L scaled = 2 1 ( y − y ^ ) 2 .
Output loss derivative ∂ L ∂ y ^ = 2 ( y ^ − y ) \frac{\partial L}{\partial \hat{y}} = 2(\hat{y} - y) ∂ y ^ ∂ L = 2 ( y ^ − y ) .
Reveal Solution
Raw Prediction Error:
e = y − y ^ = 0.0 − 0.80 = − 0.8000 e = y - \hat{y} = 0.0 - 0.80 = \mathbf{-0.8000} e = y − y ^ = 0.0 − 0.80 = − 0.8000
Standard MSE Loss:
L = ( y − y ^ ) 2 = ( 0.0 − 0.80 ) 2 = ( − 0.80 ) 2 = 0.6400 L = (y - \hat{y})^2 = (0.0 - 0.80)^2 = (-0.80)^2 = \mathbf{0.6400} L = ( y − y ^ ) 2 = ( 0.0 − 0.80 ) 2 = ( − 0.80 ) 2 = 0.6400
Scaled MSE Loss:
L scaled = 1 2 ( y − y ^ ) 2 = 1 2 ( 0.6400 ) = 0.3200 L_{\text{scaled}} = \frac{1}{2}(y - \hat{y})^2 = \frac{1}{2}(0.6400) = \mathbf{0.3200} L scaled = 2 1 ( y − y ^ ) 2 = 2 1 ( 0.6400 ) = 0.3200
Output Loss Derivative:
∂ L ∂ y ^ = 2 ( y ^ − y ) = 2 ( 0.80 − 0.0 ) = + 1.6000 \frac{\partial L}{\partial \hat{y}} = 2(\hat{y} - y) = 2(0.80 - 0.0) = \mathbf{+1.6000} ∂ y ^ ∂ L = 2 ( y ^ − y ) = 2 ( 0.80 − 0.0 ) = + 1.6000
(Interpretation: The positive derivative + 1.6000 +1.6000 + 1.6000 indicates that increasing prediction y ^ \hat{y} y ^ increases loss, signaling that the network must reduce its output to reduce error).
Problem 4: Output Layer Error Attribution & Gradient Calculation
An output neuron has loss derivative ∂ L ∂ a ( 2 ) = 1.60 \frac{\partial L}{\partial a^{(2)}} = 1.60 ∂ a ( 2 ) ∂ L = 1.60 and pre-activation z ( 2 ) = 0.0 z^{(2)} = 0.0 z ( 2 ) = 0.0 (giving σ ( 0.0 ) = 0.50 \sigma(0.0) = 0.50 σ ( 0.0 ) = 0.50 and σ ′ ( 0.0 ) = 0.25 \sigma'(0.0) = 0.25 σ ′ ( 0.0 ) = 0.25 ). The incoming hidden activation vector is a ( 1 ) = [ 2.0 3.0 ] a^{(1)} = \begin{bmatrix} 2.0 \\ 3.0 \end{bmatrix} a ( 1 ) = [ 2.0 3.0 ] .
Compute:
Output error delta δ ( 2 ) = ∂ L ∂ a ( 2 ) ⋅ σ ′ ( z ( 2 ) ) \delta^{(2)} = \frac{\partial L}{\partial a^{(2)}} \cdot \sigma'(z^{(2)}) δ ( 2 ) = ∂ a ( 2 ) ∂ L ⋅ σ ′ ( z ( 2 ) ) .
Output weight gradient vector ∂ L ∂ W ( 2 ) = δ ( 2 ) ( a ( 1 ) ) T \frac{\partial L}{\partial W^{(2)}} = \delta^{(2)} (a^{(1)})^T ∂ W ( 2 ) ∂ L = δ ( 2 ) ( a ( 1 ) ) T .
Output bias gradient ∂ L ∂ b ( 2 ) = δ ( 2 ) \frac{\partial L}{\partial b^{(2)}} = \delta^{(2)} ∂ b ( 2 ) ∂ L = δ ( 2 ) .
Reveal Solution
Output Error Delta:
δ ( 2 ) = 1.60 × 0.25 = 0.4000 \delta^{(2)} = 1.60 \times 0.25 = \mathbf{0.4000} δ ( 2 ) = 1.60 × 0.25 = 0.4000
Output Weight Gradient Vector:
∂ L ∂ W ( 2 ) = δ ( 2 ) ( a ( 1 ) ) T = 0.4000 × [ 2.0 3.0 ] = [ 0.8000 1.2000 ] \frac{\partial L}{\partial W^{(2)}} = \delta^{(2)} (a^{(1)})^T = 0.4000 \times \begin{bmatrix} 2.0 & 3.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.8000} & \mathbf{1.2000} \end{bmatrix} ∂ W ( 2 ) ∂ L = δ ( 2 ) ( a ( 1 ) ) T = 0.4000 × [ 2.0 3.0 ] = [ 0.8000 1.2000 ]
Output Bias Gradient:
∂ L ∂ b ( 2 ) = δ ( 2 ) = 0.4000 \frac{\partial L}{\partial b^{(2)}} = \delta^{(2)} = \mathbf{0.4000} ∂ b ( 2 ) ∂ L = δ ( 2 ) = 0.4000
Problem 5: Hidden Layer Backpropagation & Single Parameter Update Step
In a 2-layer network, downstream output delta is δ ( 2 ) = 0.40 \delta^{(2)} = 0.40 δ ( 2 ) = 0.40 . The second-layer weight matrix is W ( 2 ) = [ 0.50 − 0.20 ] W^{(2)} = \begin{bmatrix} 0.50 & -0.20 \end{bmatrix} W ( 2 ) = [ 0.50 − 0.20 ] .
Layer 1 has pre-activations z ( 1 ) = [ 2.0 − 1.0 ] z^{(1)} = \begin{bmatrix} 2.0 \\ -1.0 \end{bmatrix} z ( 1 ) = [ 2.0 − 1.0 ] with ReLU activations, input vector x = [ 3.0 4.0 ] x = \begin{bmatrix} 3.0 \\ 4.0 \end{bmatrix} x = [ 3.0 4.0 ] , and initial weight W 11 ( 1 ) = 1.00 W^{(1)}_{11} = 1.00 W 11 ( 1 ) = 1.00 .
Compute:
Transposed error projection ( W ( 2 ) ) T δ ( 2 ) (W^{(2)})^T \delta^{(2)} ( W ( 2 ) ) T δ ( 2 ) .
Hidden error delta vector δ ( 1 ) = ( ( W ( 2 ) ) T δ ( 2 ) ) ⊙ f ′ ( z ( 1 ) ) \delta^{(1)} = \left((W^{(2)})^T \delta^{(2)}\right) \odot f'(z^{(1)}) δ ( 1 ) = ( ( W ( 2 ) ) T δ ( 2 ) ) ⊙ f ′ ( z ( 1 ) ) .
Hidden weight gradient matrix ∂ L ∂ W ( 1 ) = δ ( 1 ) x T \frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T ∂ W ( 1 ) ∂ L = δ ( 1 ) x T .
Updated weight W 11 , new ( 1 ) W^{(1)}_{11,\text{new}} W 11 , new ( 1 ) using step size η = 0.10 \eta = 0.10 η = 0.10 .
Reveal Solution
Transposed Weight Projection:
( W ( 2 ) ) T δ ( 2 ) = [ 0.50 − 0.20 ] ( 0.40 ) = [ 0.50 × 0.40 − 0.20 × 0.40 ] = [ 0.2000 − 0.0800 ] (W^{(2)})^T \delta^{(2)} = \begin{bmatrix} 0.50 \\ -0.20 \end{bmatrix} (0.40) = \begin{bmatrix} 0.50 \times 0.40 \\ -0.20 \times 0.40 \end{bmatrix} = \begin{bmatrix} \mathbf{0.2000} \\ \mathbf{-0.0800} \end{bmatrix} ( W ( 2 ) ) T δ ( 2 ) = [ 0.50 − 0.20 ] ( 0.40 ) = [ 0.50 × 0.40 − 0.20 × 0.40 ] = [ 0.2000 − 0.0800 ]
Hidden Activation Slopes and Error Delta:
f ′ ( z ( 1 ) ) = [ ReLU ′ ( 2.0 ) ReLU ′ ( − 1.0 ) ] = [ 1.0 0.0 ] f'(z^{(1)}) = \begin{bmatrix} \text{ReLU}'(2.0) \\ \text{ReLU}'(-1.0) \end{bmatrix} = \begin{bmatrix} 1.0 \\ 0.0 \end{bmatrix} f ′ ( z ( 1 ) ) = [ ReLU ′ ( 2.0 ) ReLU ′ ( − 1.0 ) ] = [ 1.0 0.0 ]
δ ( 1 ) = [ 0.2000 − 0.0800 ] ⊙ [ 1.0 0.0 ] = [ 0.2000 0.0000 ] \delta^{(1)} = \begin{bmatrix} 0.2000 \\ -0.0800 \end{bmatrix} \odot \begin{bmatrix} 1.0 \\ 0.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.2000} \\ \mathbf{0.0000} \end{bmatrix} δ ( 1 ) = [ 0.2000 − 0.0800 ] ⊙ [ 1.0 0.0 ] = [ 0.2000 0.0000 ]
Hidden Weight Gradient Matrix:
∂ L ∂ W ( 1 ) = δ ( 1 ) x T = [ 0.2000 0.0000 ] [ 3.0 4.0 ] = [ 0.6000 0.8000 0.0000 0.0000 ] \frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T = \begin{bmatrix} 0.2000 \\ 0.0000 \end{bmatrix} \begin{bmatrix} 3.0 & 4.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.6000} & \mathbf{0.8000} \\ \mathbf{0.0000} & \mathbf{0.0000} \end{bmatrix} ∂ W ( 1 ) ∂ L = δ ( 1 ) x T = [ 0.2000 0.0000 ] [ 3.0 4.0 ] = [ 0.6000 0.0000 0.8000 0.0000 ]
Updated Weight W 11 ( 1 ) W^{(1)}_{11} W 11 ( 1 ) :
W 11 , new ( 1 ) = W 11 , old ( 1 ) − η ∂ L ∂ W 11 ( 1 ) = 1.00 − 0.10 ( 0.6000 ) = 1.00 − 0.06 = 0.9400 W^{(1)}_{11,\text{new}} = W^{(1)}_{11,\text{old}} - \eta \frac{\partial L}{\partial W^{(1)}_{11}} = 1.00 - 0.10(0.6000) = 1.00 - 0.06 = \mathbf{0.9400} W 11 , new ( 1 ) = W 11 , old ( 1 ) − η ∂ W 11 ( 1 ) ∂ L = 1.00 − 0.10 ( 0.6000 ) = 1.00 − 0.06 = 0.9400
Part 2: Applied Scenario: The Grand VC Capstone Hand Trace
An automated venture capital investment committee uses a 2 → 2 → 1 2 \to 2 \to 1 2 → 2 → 1 Multi-Layer Perceptron to predict startup success (y = 1.0 y = 1.0 y = 1.0 ) or bankruptcy (y = 0.0 y = 0.0 y = 0.0 ).
[ Input Vector x ] [ Layer 1 Hidden Space ] [ Layer 2 Decision ]
(Team Exp, Market Size) (Scalability & Moat) (Term Sheet Greenlight)
System Configuration:
Input Feature Vector (x ∈ R 2 × 1 x \in \mathbb{R}^{2 \times 1} x ∈ R 2 × 1 ):
x = [ 1.0 2.0 ] (Founding Team Experience Measurement) (Target Market Size Measurement) x = \begin{bmatrix} 1.0 \\ 2.0 \end{bmatrix} \begin{matrix} \text{(Founding Team Experience Measurement)} \\ \text{(Target Market Size Measurement)} \end{matrix} x = [ 1.0 2.0 ] (Founding Team Experience Measurement) (Target Market Size Measurement)
Initial Layer 1 Parameters (W ( 1 ) ∈ R 2 × 2 , b ( 1 ) ∈ R 2 × 1 W^{(1)} \in \mathbb{R}^{2 \times 2}, b^{(1)} \in \mathbb{R}^{2 \times 1} W ( 1 ) ∈ R 2 × 2 , b ( 1 ) ∈ R 2 × 1 ):
W ( 1 ) = [ 0.5 0.5 0.5 0.5 ] , b ( 1 ) = [ 0.0 0.0 ] W^{(1)} = \begin{bmatrix} 0.5 & 0.5 \\ 0.5 & 0.5 \end{bmatrix}, \qquad b^{(1)} = \begin{bmatrix} 0.0 \\ 0.0 \end{bmatrix} W ( 1 ) = [ 0.5 0.5 0.5 0.5 ] , b ( 1 ) = [ 0.0 0.0 ]
Initial Layer 2 Parameters (W ( 2 ) ∈ R 1 × 2 , b ( 2 ) ∈ R 1 × 1 W^{(2)} \in \mathbb{R}^{1 \times 2}, b^{(2)} \in \mathbb{R}^{1 \times 1} W ( 2 ) ∈ R 1 × 2 , b ( 2 ) ∈ R 1 × 1 ):
W ( 2 ) = [ 0.5 0.5 ] , b ( 2 ) = [ 0.0 ] W^{(2)} = \begin{bmatrix} 0.5 & 0.5 \end{bmatrix}, \qquad b^{(2)} = \begin{bmatrix} 0.0 \end{bmatrix} W ( 2 ) = [ 0.5 0.5 ] , b ( 2 ) = [ 0.0 ]
Activation Function: Standard Sigmoid σ ( z ) = 1 1 + e − z \sigma(z) = \frac{1}{1 + e^{-z}} σ ( z ) = 1 + e − z 1 across all hidden and output neurons.
Ground-Truth Observed Reality: y = 0.0 y = 0.0 y = 0.0 (The startup went bankrupt).
Step Size Factor: η = 0.50 \eta = 0.50 η = 0.50 .
Problem 6: The Grand VC Capstone Hand Trace
Execute the complete forward pass, loss calculation, backpropagation attribution, parameter updates, and second forward pass on paper:
Part A (Initial Forward Pass 1): Compute z ( 1 ) , a ( 1 ) , z ( 2 ) z^{(1)}, a^{(1)}, z^{(2)} z ( 1 ) , a ( 1 ) , z ( 2 ) , and output prediction y ^ = a ( 2 ) \hat{y} = a^{(2)} y ^ = a ( 2 ) .
Part B (Initial Loss Evaluation): Compute initial Mean Squared Error loss L 1 = ( y ^ − y ) 2 L_1 = (\hat{y} - y)^2 L 1 = ( y ^ − y ) 2 .
Part C (Output Layer Error Attribution): Compute loss derivative ∂ L ∂ a ( 2 ) \frac{\partial L}{\partial a^{(2)}} ∂ a ( 2 ) ∂ L , output activation slope σ ′ ( z ( 2 ) ) \sigma'(z^{(2)}) σ ′ ( z ( 2 ) ) , output error delta δ ( 2 ) \delta^{(2)} δ ( 2 ) , output weight gradient vector ∂ L ∂ W ( 2 ) \frac{\partial L}{\partial W^{(2)}} ∂ W ( 2 ) ∂ L , and output bias gradient ∂ L ∂ b ( 2 ) \frac{\partial L}{\partial b^{(2)}} ∂ b ( 2 ) ∂ L .
Part D (Hidden Layer Error Backpropagation): Compute backpropagated error signal ( W ( 2 ) ) T δ ( 2 ) (W^{(2)})^T \delta^{(2)} ( W ( 2 ) ) T δ ( 2 ) , hidden activation slopes σ ′ ( z ( 1 ) ) \sigma'(z^{(1)}) σ ′ ( z ( 1 ) ) , hidden error delta vector δ ( 1 ) \delta^{(1)} δ ( 1 ) , hidden weight gradient matrix ∂ L ∂ W ( 1 ) \frac{\partial L}{\partial W^{(1)}} ∂ W ( 1 ) ∂ L , and hidden bias gradient vector ∂ L ∂ b ( 1 ) \frac{\partial L}{\partial b^{(1)}} ∂ b ( 1 ) ∂ L .
Part E (Single Parameter Update Step with η = 0.50 \eta = 0.50 η = 0.50 ): Compute the full updated parameter matrices W new ( 1 ) , b new ( 1 ) , W new ( 2 ) , b new ( 2 ) W^{(1)}_{\text{new}}, b^{(1)}_{\text{new}}, W^{(2)}_{\text{new}}, b^{(2)}_{\text{new}} W new ( 1 ) , b new ( 1 ) , W new ( 2 ) , b new ( 2 ) .
Part F (Second Forward Pass 2): Re-run the complete forward pass using all updated parameters to compute z new ( 1 ) , a new ( 1 ) , z new ( 2 ) z^{(1)}_{\text{new}}, a^{(1)}_{\text{new}}, z^{(2)}_{\text{new}} z new ( 1 ) , a new ( 1 ) , z new ( 2 ) , and updated prediction y ^ new = a new ( 2 ) \hat{y}_{\text{new}} = a^{(2)}_{\text{new}} y ^ new = a new ( 2 ) .
Part G (Updated Loss & Proof of Error Reduction): Compute updated loss L 2 = ( y ^ new − y ) 2 L_2 = (\hat{y}_{\text{new}} - y)^2 L 2 = ( y ^ new − y ) 2 , calculate absolute error reduction Δ L = L 2 − L 1 \Delta L = L_2 - L_1 Δ L = L 2 − L 1 , and determine the percentage of error eliminated.
Part H (Venture Decision Engine Synthesis): In 2–3 sentences, explain why the gradient for Market Size (W 12 ( 1 ) → 0.0440 W^{(1)}_{12} \to 0.0440 W 12 ( 1 ) → 0.0440 ) was exactly twice as large as Team Experience (W 11 ( 1 ) → 0.0220 W^{(1)}_{11} \to 0.0220 W 11 ( 1 ) → 0.0220 ), and how the computational DAG automatically routes larger corrective penalties to larger input signals.
Reveal Solution
Part A: Initial Forward Pass (t = 0 t = 0 t = 0 )
Layer 1 Pre-activations (z ( 1 ) z^{(1)} z ( 1 ) ):
z 1 ( 1 ) = ( 0.5 ) ( 1.0 ) + ( 0.5 ) ( 2.0 ) + 0.0 = 0.5 + 1.0 = 1.500000 z_1^{(1)} = (0.5)(1.0) + (0.5)(2.0) + 0.0 = 0.5 + 1.0 = \mathbf{1.500000} z 1 ( 1 ) = ( 0.5 ) ( 1.0 ) + ( 0.5 ) ( 2.0 ) + 0.0 = 0.5 + 1.0 = 1.500000
z 2 ( 1 ) = ( 0.5 ) ( 1.0 ) + ( 0.5 ) ( 2.0 ) + 0.0 = 0.5 + 1.0 = 1.500000 z_2^{(1)} = (0.5)(1.0) + (0.5)(2.0) + 0.0 = 0.5 + 1.0 = \mathbf{1.500000} z 2 ( 1 ) = ( 0.5 ) ( 1.0 ) + ( 0.5 ) ( 2.0 ) + 0.0 = 0.5 + 1.0 = 1.500000
z ( 1 ) = [ 1.500000 1.500000 ] z^{(1)} = \begin{bmatrix} 1.500000 \\ 1.500000 \end{bmatrix} z ( 1 ) = [ 1.500000 1.500000 ]
Layer 1 Post-Activations (a ( 1 ) a^{(1)} a ( 1 ) ):
a 1 ( 1 ) = a 2 ( 1 ) = σ ( 1.5 ) = 1 1 + e − 1.5 = 1 1 + 0.223130 ≈ 0.817574 a_1^{(1)} = a_2^{(1)} = \sigma(1.5) = \frac{1}{1 + e^{-1.5}} = \frac{1}{1 + 0.223130} \approx \mathbf{0.817574} a 1 ( 1 ) = a 2 ( 1 ) = σ ( 1.5 ) = 1 + e − 1.5 1 = 1 + 0.223130 1 ≈ 0.817574
a ( 1 ) = [ 0.817574 0.817574 ] a^{(1)} = \begin{bmatrix} 0.817574 \\ 0.817574 \end{bmatrix} a ( 1 ) = [ 0.817574 0.817574 ]
Layer 2 Pre-activation (z ( 2 ) z^{(2)} z ( 2 ) ):
z ( 2 ) = W ( 2 ) a ( 1 ) + b ( 2 ) = ( 0.5 ) ( 0.817574 ) + ( 0.5 ) ( 0.817574 ) + 0.0 = 0.817574 z^{(2)} = W^{(2)} a^{(1)} + b^{(2)} = (0.5)(0.817574) + (0.5)(0.817574) + 0.0 = \mathbf{0.817574} z ( 2 ) = W ( 2 ) a ( 1 ) + b ( 2 ) = ( 0.5 ) ( 0.817574 ) + ( 0.5 ) ( 0.817574 ) + 0.0 = 0.817574
Layer 2 Output Prediction (y ^ = a ( 2 ) \hat{y} = a^{(2)} y ^ = a ( 2 ) ):
a ( 2 ) = σ ( 0.817574 ) = 1 1 + e − 0.817574 = 1 1 + 0.441500 ≈ 0.693721 ( 69.37 % ) a^{(2)} = \sigma(0.817574) = \frac{1}{1 + e^{-0.817574}} = \frac{1}{1 + 0.441500} \approx \mathbf{0.693721} \quad (\mathbf{69.37\%}) a ( 2 ) = σ ( 0.817574 ) = 1 + e − 0.817574 1 = 1 + 0.441500 1 ≈ 0.693721 ( 69.37% )
Part B: Initial Loss Evaluation
L 1 = ( a ( 2 ) − y ) 2 = ( 0.693721 − 0.0 ) 2 ≈ 0.481249 L_1 = (a^{(2)} - y)^2 = (0.693721 - 0.0)^2 \approx \mathbf{0.481249} L 1 = ( a ( 2 ) − y ) 2 = ( 0.693721 − 0.0 ) 2 ≈ 0.481249
Part C: Output Layer Error Attribution
Loss Derivative:
∂ L ∂ a ( 2 ) = 2 ( a ( 2 ) − y ) = 2 ( 0.693721 − 0.0 ) = 1.387442 \frac{\partial L}{\partial a^{(2)}} = 2(a^{(2)} - y) = 2(0.693721 - 0.0) = \mathbf{1.387442} ∂ a ( 2 ) ∂ L = 2 ( a ( 2 ) − y ) = 2 ( 0.693721 − 0.0 ) = 1.387442
Output Activation Slope:
σ ′ ( z ( 2 ) ) = a ( 2 ) ( 1 − a ( 2 ) ) = 0.693721 × ( 1 − 0.693721 ) = 0.693721 × 0.306279 = 0.212472 \sigma'(z^{(2)}) = a^{(2)}(1 - a^{(2)}) = 0.693721 \times (1 - 0.693721) = 0.693721 \times 0.306279 = \mathbf{0.212472} σ ′ ( z ( 2 ) ) = a ( 2 ) ( 1 − a ( 2 ) ) = 0.693721 × ( 1 − 0.693721 ) = 0.693721 × 0.306279 = 0.212472
Output Error Delta (δ ( 2 ) \delta^{(2)} δ ( 2 ) ):
δ ( 2 ) = ∂ L ∂ a ( 2 ) ⋅ σ ′ ( z ( 2 ) ) = 1.387442 × 0.212472 = 0.294793 \delta^{(2)} = \frac{\partial L}{\partial a^{(2)}} \cdot \sigma'(z^{(2)}) = 1.387442 \times 0.212472 = \mathbf{0.294793} δ ( 2 ) = ∂ a ( 2 ) ∂ L ⋅ σ ′ ( z ( 2 ) ) = 1.387442 × 0.212472 = 0.294793
Output Parameter Gradients:
∂ L ∂ W ( 2 ) = δ ( 2 ) ( a ( 1 ) ) T = 0.294793 × [ 0.817574 0.817574 ] = [ 0.241015 0.241015 ] \frac{\partial L}{\partial W^{(2)}} = \delta^{(2)} (a^{(1)})^T = 0.294793 \times \begin{bmatrix} 0.817574 & 0.817574 \end{bmatrix} = \begin{bmatrix} \mathbf{0.241015} & \mathbf{0.241015} \end{bmatrix} ∂ W ( 2 ) ∂ L = δ ( 2 ) ( a ( 1 ) ) T = 0.294793 × [ 0.817574 0.817574 ] = [ 0.241015 0.241015 ]
∂ L ∂ b ( 2 ) = δ ( 2 ) = 0.294793 \frac{\partial L}{\partial b^{(2)}} = \delta^{(2)} = \mathbf{0.294793} ∂ b ( 2 ) ∂ L = δ ( 2 ) = 0.294793
Part D: Hidden Layer Error Backpropagation
Transposed Weight Projection:
( W ( 2 ) ) T δ ( 2 ) = [ 0.5 0.5 ] ( 0.294793 ) = [ 0.147396 0.147396 ] (W^{(2)})^T \delta^{(2)} = \begin{bmatrix} 0.5 \\ 0.5 \end{bmatrix} (0.294793) = \begin{bmatrix} \mathbf{0.147396} \\ \mathbf{0.147396} \end{bmatrix} ( W ( 2 ) ) T δ ( 2 ) = [ 0.5 0.5 ] ( 0.294793 ) = [ 0.147396 0.147396 ]
Hidden Activation Slopes:
σ ′ ( z 1 ( 1 ) ) = σ ′ ( z 2 ( 1 ) ) = a 1 ( 1 ) ( 1 − a 1 ( 1 ) ) = 0.817574 × ( 1 − 0.817574 ) = 0.817574 × 0.182426 = 0.149146 \sigma'(z_1^{(1)}) = \sigma'(z_2^{(1)}) = a_1^{(1)}(1 - a_1^{(1)}) = 0.817574 \times (1 - 0.817574) = 0.817574 \times 0.182426 = \mathbf{0.149146} σ ′ ( z 1 ( 1 ) ) = σ ′ ( z 2 ( 1 ) ) = a 1 ( 1 ) ( 1 − a 1 ( 1 ) ) = 0.817574 × ( 1 − 0.817574 ) = 0.817574 × 0.182426 = 0.149146
Hidden Error Delta Vector (δ ( 1 ) \delta^{(1)} δ ( 1 ) ):
δ ( 1 ) = [ 0.147396 0.147396 ] ⊙ [ 0.149146 0.149146 ] = [ 0.021984 0.021984 ] \delta^{(1)} = \begin{bmatrix} 0.147396 \\ 0.147396 \end{bmatrix} \odot \begin{bmatrix} 0.149146 \\ 0.149146 \end{bmatrix} = \begin{bmatrix} \mathbf{0.021984} \\ \mathbf{0.021984} \end{bmatrix} δ ( 1 ) = [ 0.147396 0.147396 ] ⊙ [ 0.149146 0.149146 ] = [ 0.021984 0.021984 ]
Hidden Layer Parameter Gradients:
∂ L ∂ W ( 1 ) = δ ( 1 ) x T = [ 0.021984 0.021984 ] [ 1.0 2.0 ] = [ 0.021984 0.043967 0.021984 0.043967 ] \frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T = \begin{bmatrix} 0.021984 \\ 0.021984 \end{bmatrix} \begin{bmatrix} 1.0 & 2.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.021984} & \mathbf{0.043967} \\ \mathbf{0.021984} & \mathbf{0.043967} \end{bmatrix} ∂ W ( 1 ) ∂ L = δ ( 1 ) x T = [ 0.021984 0.021984 ] [ 1.0 2.0 ] = [ 0.021984 0.021984 0.043967 0.043967 ]
∂ L ∂ b ( 1 ) = δ ( 1 ) = [ 0.021984 0.021984 ] \frac{\partial L}{\partial b^{(1)}} = \delta^{(1)} = \begin{bmatrix} \mathbf{0.021984} \\ \mathbf{0.021984} \end{bmatrix} ∂ b ( 1 ) ∂ L = δ ( 1 ) = [ 0.021984 0.021984 ]
Part E: Single-Step Parameter Updates with η = 0.50 \eta = 0.50 η = 0.50
Layer 2 Parameters:
W new ( 2 ) = [ 0.5 0.5 ] − 0.5 [ 0.241015 0.241015 ] = [ 0.379492 0.379492 ] W^{(2)}_{\text{new}} = \begin{bmatrix} 0.5 & 0.5 \end{bmatrix} - 0.5 \begin{bmatrix} 0.241015 & 0.241015 \end{bmatrix} = \begin{bmatrix} \mathbf{0.379492} & \mathbf{0.379492} \end{bmatrix} W new ( 2 ) = [ 0.5 0.5 ] − 0.5 [ 0.241015 0.241015 ] = [ 0.379492 0.379492 ]
b new ( 2 ) = 0.0 − 0.5 ( 0.294793 ) = − 0.147396 b^{(2)}_{\text{new}} = 0.0 - 0.5(0.294793) = \mathbf{-0.147396} b new ( 2 ) = 0.0 − 0.5 ( 0.294793 ) = − 0.147396
Layer 1 Parameters:
W new ( 1 ) = [ 0.5 0.5 0.5 0.5 ] − 0.5 [ 0.021984 0.043967 0.021984 0.043967 ] = [ 0.489008 0.478016 0.489008 0.478016 ] W^{(1)}_{\text{new}} = \begin{bmatrix} 0.5 & 0.5 \\ 0.5 & 0.5 \end{bmatrix} - 0.5 \begin{bmatrix} 0.021984 & 0.043967 \\ 0.021984 & 0.043967 \end{bmatrix} = \begin{bmatrix} \mathbf{0.489008} & \mathbf{0.478016} \\ \mathbf{0.489008} & \mathbf{0.478016} \end{bmatrix} W new ( 1 ) = [ 0.5 0.5 0.5 0.5 ] − 0.5 [ 0.021984 0.021984 0.043967 0.043967 ] = [ 0.489008 0.489008 0.478016 0.478016 ]
b new ( 1 ) = [ 0.0 0.0 ] − 0.5 [ 0.021984 0.021984 ] = [ − 0.010992 − 0.010992 ] b^{(1)}_{\text{new}} = \begin{bmatrix} 0.0 \\ 0.0 \end{bmatrix} - 0.5 \begin{bmatrix} 0.021984 \\ 0.021984 \end{bmatrix} = \begin{bmatrix} \mathbf{-0.010992} \\ \mathbf{-0.010992} \end{bmatrix} b new ( 1 ) = [ 0.0 0.0 ] − 0.5 [ 0.021984 0.021984 ] = [ − 0.010992 − 0.010992 ]
Part F: Second Forward Pass (t = 1 t = 1 t = 1 )
Updated Layer 1 Pass:
z 1 , new ( 1 ) = ( 0.489008 ) ( 1.0 ) + ( 0.478016 ) ( 2.0 ) − 0.010992 = 0.489008 + 0.956033 − 0.010992 = 1.434049 z_{1, \text{new}}^{(1)} = (0.489008)(1.0) + (0.478016)(2.0) - 0.010992 = 0.489008 + 0.956033 - 0.010992 = \mathbf{1.434049} z 1 , new ( 1 ) = ( 0.489008 ) ( 1.0 ) + ( 0.478016 ) ( 2.0 ) − 0.010992 = 0.489008 + 0.956033 − 0.010992 = 1.434049
z 2 , new ( 1 ) = 1.434049 z_{2, \text{new}}^{(1)} = \mathbf{1.434049} z 2 , new ( 1 ) = 1.434049
a 1 , new ( 1 ) = a 2 , new ( 1 ) = σ ( 1.434049 ) = 1 1 + e − 1.434049 ≈ 0.807531 a_{1, \text{new}}^{(1)} = a_{2, \text{new}}^{(1)} = \sigma(1.434049) = \frac{1}{1 + e^{-1.434049}} \approx \mathbf{0.807531} a 1 , new ( 1 ) = a 2 , new ( 1 ) = σ ( 1.434049 ) = 1 + e − 1.434049 1 ≈ 0.807531
Updated Layer 2 Pass:
z new ( 2 ) = ( 0.379492 ) ( 0.807531 ) + ( 0.379492 ) ( 0.807531 ) − 0.147396 = 0.612904 − 0.147396 = 0.465508 z^{(2)}_{\text{new}} = (0.379492)(0.807531) + (0.379492)(0.807531) - 0.147396 = 0.612904 - 0.147396 = \mathbf{0.465508} z new ( 2 ) = ( 0.379492 ) ( 0.807531 ) + ( 0.379492 ) ( 0.807531 ) − 0.147396 = 0.612904 − 0.147396 = 0.465508
a new ( 2 ) = σ ( 0.465508 ) = 1 1 + e − 0.465508 ≈ 0.614320 ( 61.43 % ) a^{(2)}_{\text{new}} = \sigma(0.465508) = \frac{1}{1 + e^{-0.465508}} \approx \mathbf{0.614320} \quad (\mathbf{61.43\%}) a new ( 2 ) = σ ( 0.465508 ) = 1 + e − 0.465508 1 ≈ 0.614320 ( 61.43% )
Part G: Updated Loss & Proof of Error Reduction
Updated Loss (L 2 L_2 L 2 ):
L 2 = ( a new ( 2 ) − y ) 2 = ( 0.614320 − 0.0 ) 2 ≈ 0.377389 L_2 = (a^{(2)}_{\text{new}} - y)^2 = (0.614320 - 0.0)^2 \approx \mathbf{0.377389} L 2 = ( a new ( 2 ) − y ) 2 = ( 0.614320 − 0.0 ) 2 ≈ 0.377389
Strict Error Reduction Proof:
Δ L = L 2 − L 1 = 0.377389 − 0.481249 = − 0.103860 < 0 ( Strict Error Reduction ) \Delta L = L_2 - L_1 = 0.377389 - 0.481249 = \mathbf{-0.103860} < 0 \quad (\text{Strict Error Reduction}) Δ L = L 2 − L 1 = 0.377389 − 0.481249 = − 0.103860 < 0 ( Strict Error Reduction )
Relative Error Reduction = 0.103860 0.481249 × 100 % = 21.58 % \text{Relative Error Reduction} = \frac{0.103860}{0.481249} \times 100\% = \mathbf{21.58\%} Relative Error Reduction = 0.481249 0.103860 × 100% = 21.58%
The single parameter update step reduced prediction loss by 21.58 % 21.58\% 21.58% , lowering the false positive funding probability from 69.37 % 69.37\% 69.37% down to 61.43 % 61.43\% 61.43% .
Part H: Venture Decision Engine Synthesis
The weight gradient for Market Size (W 12 ( 1 ) → 0.0440 W^{(1)}_{12} \to 0.0440 W 12 ( 1 ) → 0.0440 ) is exactly twice as large as Team Experience (W 11 ( 1 ) → 0.0220 W^{(1)}_{11} \to 0.0220 W 11 ( 1 ) → 0.0220 ) because the input measurement for Market Size was twice as large (x 2 = 2.0 x_2 = 2.0 x 2 = 2.0 vs x 1 = 1.0 x_1 = 1.0 x 1 = 1.0 ).
Because the parameter gradient is given by ∂ L ∂ W ( 1 ) = δ ( 1 ) x T \frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T ∂ W ( 1 ) ∂ L = δ ( 1 ) x T , features with larger input magnitudes proportionally amplify the propagated error delta δ ( 1 ) \delta^{(1)} δ ( 1 ) , ensuring that backpropagation automatically assigns larger corrective parameter updates to the features that contributed most heavily to the mistaken prediction.