The Complete Neural Data Flow In Practice hero
LaboratoryThe Complete Neural Data Flow Graph

The Complete Neural Data Flow In Practice

Master complete end-to-end forward and backward passes, loss calculations, gradient attributions, and parameter updates through manual calculations.

Work through every problem on paper before expanding the step-by-step solutions.


Part 1: Abstract Mechanics

Work through these 5 rapid-fire drill problems testing tensor flow, forward passes, loss evaluation, output attribution, and hidden backpropagation.


Problem 1: End-to-End Dimension Flow Audit

Consider a 3-layer Multi-Layer Perceptron (d0→d1→d2→d3d_0 \to d_1 \to d_2 \to d_3) with layer widths d0=3d_0 = 3 (inputs), d1=4d_1 = 4 (hidden layer 1), d2=2d_2 = 2 (hidden layer 2), and d3=1d_3 = 1 (output scalar).

State the exact mathematical matrix dimensions for each of the following tensors:

  1. Input vector xx and transposed input xTx^T.
  2. Weight matrices W(1),W(2),W(3)W^{(1)}, W^{(2)}, W^{(3)} and bias vectors b(1),b(2),b(3)b^{(1)}, b^{(2)}, b^{(3)}.
  3. Pre-activation vectors z(1),z(2),z(3)z^{(1)}, z^{(2)}, z^{(3)} and activation vectors a(1),a(2),a(3)a^{(1)}, a^{(2)}, a^{(3)}.
  4. Error deltas δ(3),δ(2),δ(1)\delta^{(3)}, \delta^{(2)}, \delta^{(1)}.
  5. Parameter gradient matrices ∂L∂W(3),∂L∂W(2),∂L∂W(1)\frac{\partial L}{\partial W^{(3)}}, \frac{\partial L}{\partial W^{(2)}}, \frac{\partial L}{\partial W^{(1)}}.
Reveal Solution
  1. Input Tensors:

    • x∈R3×1x \in \mathbb{R}^{3 \times 1}
    • xT∈R1×3x^T \in \mathbb{R}^{1 \times 3}
  2. Parameter Tensors (dout×dind_{\text{out}} \times d_{\text{in}}):

    • W(1)∈R4×3,b(1)∈R4×1W^{(1)} \in \mathbb{R}^{4 \times 3}, \quad b^{(1)} \in \mathbb{R}^{4 \times 1}
    • W(2)∈R2×4,b(2)∈R2×1W^{(2)} \in \mathbb{R}^{2 \times 4}, \quad b^{(2)} \in \mathbb{R}^{2 \times 1}
    • W(3)∈R1×2,b(3)∈R1×1W^{(3)} \in \mathbb{R}^{1 \times 2}, \quad b^{(3)} \in \mathbb{R}^{1 \times 1}
  3. Forward Activation Tensors:

    • z(1),a(1)∈R4×1z^{(1)}, a^{(1)} \in \mathbb{R}^{4 \times 1}
    • z(2),a(2)∈R2×1z^{(2)}, a^{(2)} \in \mathbb{R}^{2 \times 1}
    • z(3),a(3)∈R1×1z^{(3)}, a^{(3)} \in \mathbb{R}^{1 \times 1} (scalar prediction y^\hat{y})
  4. Backward Error Deltas (δ(l)=∂L∂z(l)\delta^{(l)} = \frac{\partial L}{\partial z^{(l)}}):

    • δ(3)∈R1×1\delta^{(3)} \in \mathbb{R}^{1 \times 1}
    • δ(2)∈R2×1\delta^{(2)} \in \mathbb{R}^{2 \times 1}
    • δ(1)∈R4×1\delta^{(1)} \in \mathbb{R}^{4 \times 1}
  5. Parameter Gradients (∂L∂W(l)=δ(l)(a(l−1))T\frac{\partial L}{\partial W^{(l)}} = \delta^{(l)} (a^{(l-1)})^T):

    • ∂L∂W(3)=δ(3)(a(2))T∈R1×1×R1×2=R1×2\frac{\partial L}{\partial W^{(3)}} = \delta^{(3)} (a^{(2)})^T \in \mathbb{R}^{1 \times 1} \times \mathbb{R}^{1 \times 2} = \mathbf{\mathbb{R}^{1 \times 2}}
    • ∂L∂W(2)=δ(2)(a(1))T∈R2×1×R1×4=R2×4\frac{\partial L}{\partial W^{(2)}} = \delta^{(2)} (a^{(1)})^T \in \mathbb{R}^{2 \times 1} \times \mathbb{R}^{1 \times 4} = \mathbf{\mathbb{R}^{2 \times 4}}
    • ∂L∂W(1)=δ(1)xT∈R4×1×R1×3=R4×3\frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T \in \mathbb{R}^{4 \times 1} \times \mathbb{R}^{1 \times 3} = \mathbf{\mathbb{R}^{4 \times 3}}

Verification: Every parameter gradient matrix ∂L∂W(l)\frac{\partial L}{\partial W^{(l)}} matches the exact dimensions of its corresponding forward weight matrix W(l)W^{(l)}.


Problem 2: Complete 2→2→12 \to 2 \to 1 Forward Pass Calculation

A 2→2→12 \to 2 \to 1 Multi-Layer Perceptron has the following parameters:

  • Input vector: x=[2.01.0]x = \begin{bmatrix} 2.0 \\ 1.0 \end{bmatrix}
  • Layer 1: W(1)=[1.02.0−1.01.0],b(1)=[−1.00.5]W^{(1)} = \begin{bmatrix} 1.0 & 2.0 \\ -1.0 & 1.0 \end{bmatrix}, \quad b^{(1)} = \begin{bmatrix} -1.0 \\ 0.5 \end{bmatrix}, with ReLU activation.
  • Layer 2: W(2)=[2.01.0],b(2)=[−1.0]W^{(2)} = \begin{bmatrix} 2.0 & 1.0 \end{bmatrix}, \quad b^{(2)} = \begin{bmatrix} -1.0 \end{bmatrix}, with Sigmoid activation.

Compute:

  1. Layer 1 pre-activation vector z(1)z^{(1)} and hidden activation vector a(1)a^{(1)}.
  2. Layer 2 pre-activation scalar z(2)z^{(2)} and output prediction y^=a(2)\hat{y} = a^{(2)} (use e−5.0≈0.006738e^{-5.0} \approx 0.006738).
Reveal Solution
  1. Layer 1 Forward Calculation:

    z1(1)=(1.0)(2.0)+(2.0)(1.0)−1.0=2.0+2.0−1.0=+3.0z_1^{(1)} = (1.0)(2.0) + (2.0)(1.0) - 1.0 = 2.0 + 2.0 - 1.0 = \mathbf{+3.0} z2(1)=(−1.0)(2.0)+(1.0)(1.0)+0.5=−2.0+1.0+0.5=−0.5z_2^{(1)} = (-1.0)(2.0) + (1.0)(1.0) + 0.5 = -2.0 + 1.0 + 0.5 = \mathbf{-0.5} z(1)=[+3.0−0.5]z^{(1)} = \begin{bmatrix} +3.0 \\ -0.5 \end{bmatrix}

    Apply ReLU:

    a(1)=[max⁡(0,3.0)max⁡(0,−0.5)]=[3.00.0]a^{(1)} = \begin{bmatrix} \max(0, 3.0) \\ \max(0, -0.5) \end{bmatrix} = \begin{bmatrix} \mathbf{3.0} \\ \mathbf{0.0} \end{bmatrix}
  2. Layer 2 Forward Calculation:

    z(2)=W(2)a(1)+b(2)=(2.0)(3.0)+(1.0)(0.0)−1.0=6.0+0.0−1.0=+5.0z^{(2)} = W^{(2)} a^{(1)} + b^{(2)} = (2.0)(3.0) + (1.0)(0.0) - 1.0 = 6.0 + 0.0 - 1.0 = \mathbf{+5.0}

    Apply Sigmoid:

    a(2)=σ(5.0)=11+e−5.0=11+0.006738=11.006738≈0.993307(99.33%)a^{(2)} = \sigma(5.0) = \frac{1}{1 + e^{-5.0}} = \frac{1}{1 + 0.006738} = \frac{1}{1.006738} \approx \mathbf{0.993307} \quad (\mathbf{99.33\%})

Problem 3: Mean Squared Error Loss & Sensitivity

A neural network outputs prediction y^=0.80\hat{y} = 0.80 for a training instance where ground truth is y=0.0y = 0.0.

Compute:

  1. Raw prediction error e=y−y^e = y - \hat{y}.
  2. Standard Mean Squared Error loss L=(y−y^)2L = (y - \hat{y})^2.
  3. Scaled MSE loss Lscaled=12(y−y^)2L_{\text{scaled}} = \frac{1}{2}(y - \hat{y})^2.
  4. Output loss derivative ∂L∂y^=2(y^−y)\frac{\partial L}{\partial \hat{y}} = 2(\hat{y} - y).
Reveal Solution
  1. Raw Prediction Error:

    e=y−y^=0.0−0.80=−0.8000e = y - \hat{y} = 0.0 - 0.80 = \mathbf{-0.8000}
  2. Standard MSE Loss:

    L=(y−y^)2=(0.0−0.80)2=(−0.80)2=0.6400L = (y - \hat{y})^2 = (0.0 - 0.80)^2 = (-0.80)^2 = \mathbf{0.6400}
  3. Scaled MSE Loss:

    Lscaled=12(y−y^)2=12(0.6400)=0.3200L_{\text{scaled}} = \frac{1}{2}(y - \hat{y})^2 = \frac{1}{2}(0.6400) = \mathbf{0.3200}
  4. Output Loss Derivative:

    ∂L∂y^=2(y^−y)=2(0.80−0.0)=+1.6000\frac{\partial L}{\partial \hat{y}} = 2(\hat{y} - y) = 2(0.80 - 0.0) = \mathbf{+1.6000}

(Interpretation: The positive derivative +1.6000+1.6000 indicates that increasing prediction y^\hat{y} increases loss, signaling that the network must reduce its output to reduce error).


Problem 4: Output Layer Error Attribution & Gradient Calculation

An output neuron has loss derivative ∂L∂a(2)=1.60\frac{\partial L}{\partial a^{(2)}} = 1.60 and pre-activation z(2)=0.0z^{(2)} = 0.0 (giving σ(0.0)=0.50\sigma(0.0) = 0.50 and σ′(0.0)=0.25\sigma'(0.0) = 0.25). The incoming hidden activation vector is a(1)=[2.03.0]a^{(1)} = \begin{bmatrix} 2.0 \\ 3.0 \end{bmatrix}.

Compute:

  1. Output error delta δ(2)=∂L∂a(2)⋅σ′(z(2))\delta^{(2)} = \frac{\partial L}{\partial a^{(2)}} \cdot \sigma'(z^{(2)}).
  2. Output weight gradient vector ∂L∂W(2)=δ(2)(a(1))T\frac{\partial L}{\partial W^{(2)}} = \delta^{(2)} (a^{(1)})^T.
  3. Output bias gradient ∂L∂b(2)=δ(2)\frac{\partial L}{\partial b^{(2)}} = \delta^{(2)}.
Reveal Solution
  1. Output Error Delta:

    δ(2)=1.60×0.25=0.4000\delta^{(2)} = 1.60 \times 0.25 = \mathbf{0.4000}
  2. Output Weight Gradient Vector:

    ∂L∂W(2)=δ(2)(a(1))T=0.4000×[2.03.0]=[0.80001.2000]\frac{\partial L}{\partial W^{(2)}} = \delta^{(2)} (a^{(1)})^T = 0.4000 \times \begin{bmatrix} 2.0 & 3.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.8000} & \mathbf{1.2000} \end{bmatrix}
  3. Output Bias Gradient:

    ∂L∂b(2)=δ(2)=0.4000\frac{\partial L}{\partial b^{(2)}} = \delta^{(2)} = \mathbf{0.4000}

Problem 5: Hidden Layer Backpropagation & Single Parameter Update Step

In a 2-layer network, downstream output delta is δ(2)=0.40\delta^{(2)} = 0.40. The second-layer weight matrix is W(2)=[0.50−0.20]W^{(2)} = \begin{bmatrix} 0.50 & -0.20 \end{bmatrix}.

Layer 1 has pre-activations z(1)=[2.0−1.0]z^{(1)} = \begin{bmatrix} 2.0 \\ -1.0 \end{bmatrix} with ReLU activations, input vector x=[3.04.0]x = \begin{bmatrix} 3.0 \\ 4.0 \end{bmatrix}, and initial weight W11(1)=1.00W^{(1)}_{11} = 1.00.

Compute:

  1. Transposed error projection (W(2))Tδ(2)(W^{(2)})^T \delta^{(2)}.
  2. Hidden error delta vector δ(1)=((W(2))Tδ(2))⊙f′(z(1))\delta^{(1)} = \left((W^{(2)})^T \delta^{(2)}\right) \odot f'(z^{(1)}).
  3. Hidden weight gradient matrix ∂L∂W(1)=δ(1)xT\frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T.
  4. Updated weight W11,new(1)W^{(1)}_{11,\text{new}} using step size η=0.10\eta = 0.10.
Reveal Solution
  1. Transposed Weight Projection:

    (W(2))Tδ(2)=[0.50−0.20](0.40)=[0.50×0.40−0.20×0.40]=[0.2000−0.0800](W^{(2)})^T \delta^{(2)} = \begin{bmatrix} 0.50 \\ -0.20 \end{bmatrix} (0.40) = \begin{bmatrix} 0.50 \times 0.40 \\ -0.20 \times 0.40 \end{bmatrix} = \begin{bmatrix} \mathbf{0.2000} \\ \mathbf{-0.0800} \end{bmatrix}
  2. Hidden Activation Slopes and Error Delta:

    f′(z(1))=[ReLU′(2.0)ReLU′(−1.0)]=[1.00.0]f'(z^{(1)}) = \begin{bmatrix} \text{ReLU}'(2.0) \\ \text{ReLU}'(-1.0) \end{bmatrix} = \begin{bmatrix} 1.0 \\ 0.0 \end{bmatrix} δ(1)=[0.2000−0.0800]⊙[1.00.0]=[0.20000.0000]\delta^{(1)} = \begin{bmatrix} 0.2000 \\ -0.0800 \end{bmatrix} \odot \begin{bmatrix} 1.0 \\ 0.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.2000} \\ \mathbf{0.0000} \end{bmatrix}
  3. Hidden Weight Gradient Matrix:

    ∂L∂W(1)=δ(1)xT=[0.20000.0000][3.04.0]=[0.60000.80000.00000.0000]\frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T = \begin{bmatrix} 0.2000 \\ 0.0000 \end{bmatrix} \begin{bmatrix} 3.0 & 4.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.6000} & \mathbf{0.8000} \\ \mathbf{0.0000} & \mathbf{0.0000} \end{bmatrix}
  4. Updated Weight W11(1)W^{(1)}_{11}:

    W11,new(1)=W11,old(1)−η∂L∂W11(1)=1.00−0.10(0.6000)=1.00−0.06=0.9400W^{(1)}_{11,\text{new}} = W^{(1)}_{11,\text{old}} - \eta \frac{\partial L}{\partial W^{(1)}_{11}} = 1.00 - 0.10(0.6000) = 1.00 - 0.06 = \mathbf{0.9400}

Part 2: Applied Scenario: The Grand VC Capstone Hand Trace

An automated venture capital investment committee uses a 2→2→12 \to 2 \to 1 Multi-Layer Perceptron to predict startup success (y=1.0y = 1.0) or bankruptcy (y=0.0y = 0.0).

[ Input Vector x ]          [ Layer 1 Hidden Space ]          [ Layer 2 Decision ]
(Team Exp, Market Size)     (Scalability & Moat)             (Term Sheet Greenlight)

System Configuration:

  • Input Feature Vector (x∈R2×1x \in \mathbb{R}^{2 \times 1}):

    x=[1.02.0](Founding Team Experience Measurement)(Target Market Size Measurement)x = \begin{bmatrix} 1.0 \\ 2.0 \end{bmatrix} \begin{matrix} \text{(Founding Team Experience Measurement)} \\ \text{(Target Market Size Measurement)} \end{matrix}
  • Initial Layer 1 Parameters (W(1)∈R2×2,b(1)∈R2×1W^{(1)} \in \mathbb{R}^{2 \times 2}, b^{(1)} \in \mathbb{R}^{2 \times 1}):

    W(1)=[0.50.50.50.5],b(1)=[0.00.0]W^{(1)} = \begin{bmatrix} 0.5 & 0.5 \\ 0.5 & 0.5 \end{bmatrix}, \qquad b^{(1)} = \begin{bmatrix} 0.0 \\ 0.0 \end{bmatrix}
  • Initial Layer 2 Parameters (W(2)∈R1×2,b(2)∈R1×1W^{(2)} \in \mathbb{R}^{1 \times 2}, b^{(2)} \in \mathbb{R}^{1 \times 1}):

    W(2)=[0.50.5],b(2)=[0.0]W^{(2)} = \begin{bmatrix} 0.5 & 0.5 \end{bmatrix}, \qquad b^{(2)} = \begin{bmatrix} 0.0 \end{bmatrix}
  • Activation Function: Standard Sigmoid σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}} across all hidden and output neurons.

  • Ground-Truth Observed Reality: y=0.0y = 0.0 (The startup went bankrupt).

  • Step Size Factor: η=0.50\eta = 0.50.


Problem 6: The Grand VC Capstone Hand Trace

Execute the complete forward pass, loss calculation, backpropagation attribution, parameter updates, and second forward pass on paper:

  • Part A (Initial Forward Pass 1): Compute z(1),a(1),z(2)z^{(1)}, a^{(1)}, z^{(2)}, and output prediction y^=a(2)\hat{y} = a^{(2)}.
  • Part B (Initial Loss Evaluation): Compute initial Mean Squared Error loss L1=(y^−y)2L_1 = (\hat{y} - y)^2.
  • Part C (Output Layer Error Attribution): Compute loss derivative ∂L∂a(2)\frac{\partial L}{\partial a^{(2)}}, output activation slope σ′(z(2))\sigma'(z^{(2)}), output error delta δ(2)\delta^{(2)}, output weight gradient vector ∂L∂W(2)\frac{\partial L}{\partial W^{(2)}}, and output bias gradient ∂L∂b(2)\frac{\partial L}{\partial b^{(2)}}.
  • Part D (Hidden Layer Error Backpropagation): Compute backpropagated error signal (W(2))Tδ(2)(W^{(2)})^T \delta^{(2)}, hidden activation slopes σ′(z(1))\sigma'(z^{(1)}), hidden error delta vector δ(1)\delta^{(1)}, hidden weight gradient matrix ∂L∂W(1)\frac{\partial L}{\partial W^{(1)}}, and hidden bias gradient vector ∂L∂b(1)\frac{\partial L}{\partial b^{(1)}}.
  • Part E (Single Parameter Update Step with η=0.50\eta = 0.50): Compute the full updated parameter matrices Wnew(1),bnew(1),Wnew(2),bnew(2)W^{(1)}_{\text{new}}, b^{(1)}_{\text{new}}, W^{(2)}_{\text{new}}, b^{(2)}_{\text{new}}.
  • Part F (Second Forward Pass 2): Re-run the complete forward pass using all updated parameters to compute znew(1),anew(1),znew(2)z^{(1)}_{\text{new}}, a^{(1)}_{\text{new}}, z^{(2)}_{\text{new}}, and updated prediction y^new=anew(2)\hat{y}_{\text{new}} = a^{(2)}_{\text{new}}.
  • Part G (Updated Loss & Proof of Error Reduction): Compute updated loss L2=(y^new−y)2L_2 = (\hat{y}_{\text{new}} - y)^2, calculate absolute error reduction ΔL=L2−L1\Delta L = L_2 - L_1, and determine the percentage of error eliminated.
  • Part H (Venture Decision Engine Synthesis): In 2–3 sentences, explain why the gradient for Market Size (W12(1)→0.0440W^{(1)}_{12} \to 0.0440) was exactly twice as large as Team Experience (W11(1)→0.0220W^{(1)}_{11} \to 0.0220), and how the computational DAG automatically routes larger corrective penalties to larger input signals.
Reveal Solution

Part A: Initial Forward Pass (t=0t = 0)

  1. Layer 1 Pre-activations (z(1)z^{(1)}):

    z1(1)=(0.5)(1.0)+(0.5)(2.0)+0.0=0.5+1.0=1.500000z_1^{(1)} = (0.5)(1.0) + (0.5)(2.0) + 0.0 = 0.5 + 1.0 = \mathbf{1.500000} z2(1)=(0.5)(1.0)+(0.5)(2.0)+0.0=0.5+1.0=1.500000z_2^{(1)} = (0.5)(1.0) + (0.5)(2.0) + 0.0 = 0.5 + 1.0 = \mathbf{1.500000} z(1)=[1.5000001.500000]z^{(1)} = \begin{bmatrix} 1.500000 \\ 1.500000 \end{bmatrix}
  2. Layer 1 Post-Activations (a(1)a^{(1)}):

    a1(1)=a2(1)=σ(1.5)=11+e−1.5=11+0.223130≈0.817574a_1^{(1)} = a_2^{(1)} = \sigma(1.5) = \frac{1}{1 + e^{-1.5}} = \frac{1}{1 + 0.223130} \approx \mathbf{0.817574} a(1)=[0.8175740.817574]a^{(1)} = \begin{bmatrix} 0.817574 \\ 0.817574 \end{bmatrix}
  3. Layer 2 Pre-activation (z(2)z^{(2)}):

    z(2)=W(2)a(1)+b(2)=(0.5)(0.817574)+(0.5)(0.817574)+0.0=0.817574z^{(2)} = W^{(2)} a^{(1)} + b^{(2)} = (0.5)(0.817574) + (0.5)(0.817574) + 0.0 = \mathbf{0.817574}
  4. Layer 2 Output Prediction (y^=a(2)\hat{y} = a^{(2)}):

    a(2)=σ(0.817574)=11+e−0.817574=11+0.441500≈0.693721(69.37%)a^{(2)} = \sigma(0.817574) = \frac{1}{1 + e^{-0.817574}} = \frac{1}{1 + 0.441500} \approx \mathbf{0.693721} \quad (\mathbf{69.37\%})

Part B: Initial Loss Evaluation

L1=(a(2)−y)2=(0.693721−0.0)2≈0.481249L_1 = (a^{(2)} - y)^2 = (0.693721 - 0.0)^2 \approx \mathbf{0.481249}

Part C: Output Layer Error Attribution

  1. Loss Derivative:

    ∂L∂a(2)=2(a(2)−y)=2(0.693721−0.0)=1.387442\frac{\partial L}{\partial a^{(2)}} = 2(a^{(2)} - y) = 2(0.693721 - 0.0) = \mathbf{1.387442}
  2. Output Activation Slope:

    σ′(z(2))=a(2)(1−a(2))=0.693721×(1−0.693721)=0.693721×0.306279=0.212472\sigma'(z^{(2)}) = a^{(2)}(1 - a^{(2)}) = 0.693721 \times (1 - 0.693721) = 0.693721 \times 0.306279 = \mathbf{0.212472}
  3. Output Error Delta (δ(2)\delta^{(2)}):

    δ(2)=∂L∂a(2)⋅σ′(z(2))=1.387442×0.212472=0.294793\delta^{(2)} = \frac{\partial L}{\partial a^{(2)}} \cdot \sigma'(z^{(2)}) = 1.387442 \times 0.212472 = \mathbf{0.294793}
  4. Output Parameter Gradients:

    ∂L∂W(2)=δ(2)(a(1))T=0.294793×[0.8175740.817574]=[0.2410150.241015]\frac{\partial L}{\partial W^{(2)}} = \delta^{(2)} (a^{(1)})^T = 0.294793 \times \begin{bmatrix} 0.817574 & 0.817574 \end{bmatrix} = \begin{bmatrix} \mathbf{0.241015} & \mathbf{0.241015} \end{bmatrix} ∂L∂b(2)=δ(2)=0.294793\frac{\partial L}{\partial b^{(2)}} = \delta^{(2)} = \mathbf{0.294793}

Part D: Hidden Layer Error Backpropagation

  1. Transposed Weight Projection:

    (W(2))Tδ(2)=[0.50.5](0.294793)=[0.1473960.147396](W^{(2)})^T \delta^{(2)} = \begin{bmatrix} 0.5 \\ 0.5 \end{bmatrix} (0.294793) = \begin{bmatrix} \mathbf{0.147396} \\ \mathbf{0.147396} \end{bmatrix}
  2. Hidden Activation Slopes:

    σ′(z1(1))=σ′(z2(1))=a1(1)(1−a1(1))=0.817574×(1−0.817574)=0.817574×0.182426=0.149146\sigma'(z_1^{(1)}) = \sigma'(z_2^{(1)}) = a_1^{(1)}(1 - a_1^{(1)}) = 0.817574 \times (1 - 0.817574) = 0.817574 \times 0.182426 = \mathbf{0.149146}
  3. Hidden Error Delta Vector (δ(1)\delta^{(1)}):

    δ(1)=[0.1473960.147396]⊙[0.1491460.149146]=[0.0219840.021984]\delta^{(1)} = \begin{bmatrix} 0.147396 \\ 0.147396 \end{bmatrix} \odot \begin{bmatrix} 0.149146 \\ 0.149146 \end{bmatrix} = \begin{bmatrix} \mathbf{0.021984} \\ \mathbf{0.021984} \end{bmatrix}
  4. Hidden Layer Parameter Gradients:

    ∂L∂W(1)=δ(1)xT=[0.0219840.021984][1.02.0]=[0.0219840.0439670.0219840.043967]\frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T = \begin{bmatrix} 0.021984 \\ 0.021984 \end{bmatrix} \begin{bmatrix} 1.0 & 2.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.021984} & \mathbf{0.043967} \\ \mathbf{0.021984} & \mathbf{0.043967} \end{bmatrix} ∂L∂b(1)=δ(1)=[0.0219840.021984]\frac{\partial L}{\partial b^{(1)}} = \delta^{(1)} = \begin{bmatrix} \mathbf{0.021984} \\ \mathbf{0.021984} \end{bmatrix}

Part E: Single-Step Parameter Updates with η=0.50\eta = 0.50

  1. Layer 2 Parameters:

    Wnew(2)=[0.50.5]−0.5[0.2410150.241015]=[0.3794920.379492]W^{(2)}_{\text{new}} = \begin{bmatrix} 0.5 & 0.5 \end{bmatrix} - 0.5 \begin{bmatrix} 0.241015 & 0.241015 \end{bmatrix} = \begin{bmatrix} \mathbf{0.379492} & \mathbf{0.379492} \end{bmatrix} bnew(2)=0.0−0.5(0.294793)=−0.147396b^{(2)}_{\text{new}} = 0.0 - 0.5(0.294793) = \mathbf{-0.147396}
  2. Layer 1 Parameters:

    Wnew(1)=[0.50.50.50.5]−0.5[0.0219840.0439670.0219840.043967]=[0.4890080.4780160.4890080.478016]W^{(1)}_{\text{new}} = \begin{bmatrix} 0.5 & 0.5 \\ 0.5 & 0.5 \end{bmatrix} - 0.5 \begin{bmatrix} 0.021984 & 0.043967 \\ 0.021984 & 0.043967 \end{bmatrix} = \begin{bmatrix} \mathbf{0.489008} & \mathbf{0.478016} \\ \mathbf{0.489008} & \mathbf{0.478016} \end{bmatrix} bnew(1)=[0.00.0]−0.5[0.0219840.021984]=[−0.010992−0.010992]b^{(1)}_{\text{new}} = \begin{bmatrix} 0.0 \\ 0.0 \end{bmatrix} - 0.5 \begin{bmatrix} 0.021984 \\ 0.021984 \end{bmatrix} = \begin{bmatrix} \mathbf{-0.010992} \\ \mathbf{-0.010992} \end{bmatrix}

Part F: Second Forward Pass (t=1t = 1)

  1. Updated Layer 1 Pass:

    z1,new(1)=(0.489008)(1.0)+(0.478016)(2.0)−0.010992=0.489008+0.956033−0.010992=1.434049z_{1, \text{new}}^{(1)} = (0.489008)(1.0) + (0.478016)(2.0) - 0.010992 = 0.489008 + 0.956033 - 0.010992 = \mathbf{1.434049} z2,new(1)=1.434049z_{2, \text{new}}^{(1)} = \mathbf{1.434049} a1,new(1)=a2,new(1)=σ(1.434049)=11+e−1.434049≈0.807531a_{1, \text{new}}^{(1)} = a_{2, \text{new}}^{(1)} = \sigma(1.434049) = \frac{1}{1 + e^{-1.434049}} \approx \mathbf{0.807531}
  2. Updated Layer 2 Pass:

    znew(2)=(0.379492)(0.807531)+(0.379492)(0.807531)−0.147396=0.612904−0.147396=0.465508z^{(2)}_{\text{new}} = (0.379492)(0.807531) + (0.379492)(0.807531) - 0.147396 = 0.612904 - 0.147396 = \mathbf{0.465508} anew(2)=σ(0.465508)=11+e−0.465508≈0.614320(61.43%)a^{(2)}_{\text{new}} = \sigma(0.465508) = \frac{1}{1 + e^{-0.465508}} \approx \mathbf{0.614320} \quad (\mathbf{61.43\%})

Part G: Updated Loss & Proof of Error Reduction

  1. Updated Loss (L2L_2):

    L2=(anew(2)−y)2=(0.614320−0.0)2≈0.377389L_2 = (a^{(2)}_{\text{new}} - y)^2 = (0.614320 - 0.0)^2 \approx \mathbf{0.377389}
  2. Strict Error Reduction Proof:

    ΔL=L2−L1=0.377389−0.481249=−0.103860<0(Strict Error Reduction)\Delta L = L_2 - L_1 = 0.377389 - 0.481249 = \mathbf{-0.103860} < 0 \quad (\text{Strict Error Reduction}) Relative Error Reduction=0.1038600.481249×100%=21.58%\text{Relative Error Reduction} = \frac{0.103860}{0.481249} \times 100\% = \mathbf{21.58\%}

The single parameter update step reduced prediction loss by 21.58%21.58\%, lowering the false positive funding probability from 69.37%69.37\% down to 61.43%61.43\%.


Part H: Venture Decision Engine Synthesis

The weight gradient for Market Size (W12(1)→0.0440W^{(1)}_{12} \to 0.0440) is exactly twice as large as Team Experience (W11(1)→0.0220W^{(1)}_{11} \to 0.0220) because the input measurement for Market Size was twice as large (x2=2.0x_2 = 2.0 vs x1=1.0x_1 = 1.0).

Because the parameter gradient is given by ∂L∂W(1)=δ(1)xT\frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T, features with larger input magnitudes proportionally amplify the propagated error delta δ(1)\delta^{(1)}, ensuring that backpropagation automatically assigns larger corrective parameter updates to the features that contributed most heavily to the mistaken prediction.

Previous
Operating the Grand Capstone Engine