The Closed Learning Cycle on Paper hero
Lesson 2The Complete Neural Data Flow Graph

The Closed Learning Cycle on Paper

Execute a complete forward pass, loss calculation, backpropagation cycle, and second forward pass on paper to prove error reduction.

Let us trace the complete, unbroken learning cycle on paper with explicit numbers for our running 2-layer network.


System Configuration: 'Die Hard in Space'

We evaluate the screenplay for 'Die Hard in Space' across our running 4→2→14 \to 2 \to 1 Multi-Layer Perceptron.

  • Input Feature Vector (x∈R4×1x \in \mathbb{R}^{4 \times 1}):

    x=[1.00.00.51.0](Action Measurement)(Romance Measurement)(Comedy Measurement)(Sci-Fi Measurement)x = \begin{bmatrix} 1.0 \\ 0.0 \\ 0.5 \\ 1.0 \end{bmatrix} \begin{matrix} \text{(Action Measurement)} \\ \text{(Romance Measurement)} \\ \text{(Comedy Measurement)} \\ \text{(Sci-Fi Measurement)} \end{matrix}
  • Layer 1 Initial Parameters (W(1)∈R2×4,b(1)∈R2×1W^{(1)} \in \mathbb{R}^{2 \times 4}, b^{(1)} \in \mathbb{R}^{2 \times 1}):

    W(1)=[3.0−2.00.03.0−2.03.02.0−1.0],b(1)=[−2.0−1.5]W^{(1)} = \begin{bmatrix} 3.0 & -2.0 & 0.0 & 3.0 \\ -2.0 & 3.0 & 2.0 & -1.0 \end{bmatrix}, \qquad b^{(1)} = \begin{bmatrix} -2.0 \\ -1.5 \end{bmatrix}

    Hidden activation function: f(z)=ReLU(z)=max⁡(0,z)f(z) = \text{ReLU}(z) = \max(0, z).

  • Layer 2 Initial Parameters (W(2)∈R1×2,b(2)∈R1×1W^{(2)} \in \mathbb{R}^{1 \times 2}, b^{(2)} \in \mathbb{R}^{1 \times 1}):

    W(2)=[1.51.0],b(2)=[−2.0]W^{(2)} = \begin{bmatrix} 1.5 & 1.0 \end{bmatrix}, \qquad b^{(2)} = \begin{bmatrix} -2.0 \end{bmatrix}

    Output activation function: σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}.

  • Ground-Truth Target: y=0.0y = 0.0 (Box Office Flop).

  • Step Size Factor: η=0.50\eta = 0.50.


Phase 1: Initial Forward Pass (t=0t = 0)

Step 1: Layer 1 Affine Transformation

Compute z(1)=W(1)x+b(1)z^{(1)} = W^{(1)} x + b^{(1)}:

z1(1)=(3.0)(1.0)+(−2.0)(0.0)+(0.0)(0.5)+(3.0)(1.0)−2.0=3.0+0.0+0.0+3.0−2.0=+4.0z_1^{(1)} = (3.0)(1.0) + (-2.0)(0.0) + (0.0)(0.5) + (3.0)(1.0) - 2.0 = 3.0 + 0.0 + 0.0 + 3.0 - 2.0 = \mathbf{+4.0} z2(1)=(−2.0)(1.0)+(3.0)(0.0)+(2.0)(0.5)+(−1.0)(1.0)−1.5=−2.0+0.0+1.0−1.0−1.5=−3.5z_2^{(1)} = (-2.0)(1.0) + (3.0)(0.0) + (2.0)(0.5) + (-1.0)(1.0) - 1.5 = -2.0 + 0.0 + 1.0 - 1.0 - 1.5 = \mathbf{-3.5} z(1)=[+4.0−3.5]z^{(1)} = \begin{bmatrix} +4.0 \\ -3.5 \end{bmatrix}

Step 2: Layer 1 Non-Linear Activation

Apply element-wise ReLU activation:

a(1)=ReLU(z(1))=[max⁡(0,4.0)max⁡(0,−3.5)]=[4.00.0](Summer Popcorn Flick Factor)(Rom-Com Factor)a^{(1)} = \text{ReLU}\left(z^{(1)}\right) = \begin{bmatrix} \max(0, 4.0) \\ \max(0, -3.5) \end{bmatrix} = \begin{bmatrix} \mathbf{4.0} \\ \mathbf{0.0} \end{bmatrix} \begin{matrix} \text{(Summer Popcorn Flick Factor)} \\ \text{(Rom-Com Factor)} \end{matrix}

Step 3: Layer 2 Affine Transformation

Compute scalar pre-activation z(2)=W(2)a(1)+b(2)z^{(2)} = W^{(2)} a^{(1)} + b^{(2)}:

z(2)=(1.5)(4.0)+(1.0)(0.0)+(−2.0)=6.0+0.0−2.0=+4.0z^{(2)} = (1.5)(4.0) + (1.0)(0.0) + (-2.0) = 6.0 + 0.0 - 2.0 = \mathbf{+4.0}

Step 4: Layer 2 Sigmoid Probability Squashing

Compute final prediction y^=a(2)=σ(z(2))\hat{y} = a^{(2)} = \sigma(z^{(2)}):

a(2)=σ(4.0)=11+e−4.0=11+0.018316≈0.982014(98.20%)a^{(2)} = \sigma(4.0) = \frac{1}{1 + e^{-4.0}} = \frac{1}{1 + 0.018316} \approx \mathbf{0.982014} \quad (\mathbf{98.20\%})

Phase 2: Error Quantification & Loss Evaluation

The network predicted a 98.20%98.20\% Hit probability, but the movie was a box office Flop (y=0.0y = 0.0).

Step 1: Raw Prediction Error

e=y−y^=0.0−0.982014=−0.982014e = y - \hat{y} = 0.0 - 0.982014 = \mathbf{-0.982014}

Step 2: Mean Squared Error Loss

L1=(a(2)−y)2=(0.982014−0.0)2≈0.964351L_1 = (a^{(2)} - y)^2 = (0.982014 - 0.0)^2 \approx \mathbf{0.964351}

Phase 3: Backward Pass & Gradient Attribution

Step 1: Output Layer Error Attribution

  1. Loss Sensitivity:

    ∂L∂a(2)=2(a(2)−y)=2(0.982014−0.0)=+1.964028\frac{\partial L}{\partial a^{(2)}} = 2(a^{(2)} - y) = 2(0.982014 - 0.0) = \mathbf{+1.964028}
  2. Activation Function Slope:

    σ′(z(2))=a(2)(1−a(2))=0.982014×(1−0.982014)=0.982014×0.017986=0.017663\sigma'(z^{(2)}) = a^{(2)}(1 - a^{(2)}) = 0.982014 \times (1 - 0.982014) = 0.982014 \times 0.017986 = \mathbf{0.017663}
  3. Output Error Delta (δ(2)\delta^{(2)}):

    δ(2)=∂L∂a(2)⋅σ′(z(2))=(+1.964028)×(0.017663)=+0.034690\delta^{(2)} = \frac{\partial L}{\partial a^{(2)}} \cdot \sigma'(z^{(2)}) = (+1.964028) \times (0.017663) = \mathbf{+0.034690}
  4. Output Parameter Gradients:

    ∂L∂W(2)=δ(2)(a(1))T=0.034690×[4.00.0]=[0.1387600.000000]\frac{\partial L}{\partial W^{(2)}} = \delta^{(2)} (a^{(1)})^T = 0.034690 \times \begin{bmatrix} 4.0 & 0.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.138760} & \mathbf{0.000000} \end{bmatrix} ∂L∂b(2)=δ(2)=+0.034690\frac{\partial L}{\partial b^{(2)}} = \delta^{(2)} = \mathbf{+0.034690}

Step 2: Hidden Layer Error Backpropagation

  1. Transposed Weight Projection:

    (W(2))Tδ(2)=[1.51.0](0.034690)=[1.5×0.0346901.0×0.034690]=[0.0520350.034690](W^{(2)})^T \delta^{(2)} = \begin{bmatrix} 1.5 \\ 1.0 \end{bmatrix} (0.034690) = \begin{bmatrix} 1.5 \times 0.034690 \\ 1.0 \times 0.034690 \end{bmatrix} = \begin{bmatrix} \mathbf{0.052035} \\ \mathbf{0.034690} \end{bmatrix}
  2. Hidden Activation Slopes:

    f′(z(1))=[ReLU′(+4.0)ReLU′(−3.5)]=[1.00.0]f'(z^{(1)}) = \begin{bmatrix} \text{ReLU}'(+4.0) \\ \text{ReLU}'(-3.5) \end{bmatrix} = \begin{bmatrix} \mathbf{1.0} \\ \mathbf{0.0} \end{bmatrix}
  3. Hidden Error Delta (δ(1)\delta^{(1)}):

    δ(1)=((W(2))Tδ(2))⊙f′(z(1))=[0.0520350.034690]⊙[1.00.0]=[0.0520350.000000]\delta^{(1)} = \left( (W^{(2)})^T \delta^{(2)} \right) \odot f'(z^{(1)}) = \begin{bmatrix} 0.052035 \\ 0.034690 \end{bmatrix} \odot \begin{bmatrix} 1.0 \\ 0.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.052035} \\ \mathbf{0.000000} \end{bmatrix}
  4. Hidden Layer Parameter Gradients:

    ∂L∂W(1)=δ(1)xT=[0.0520350.000000][1.00.00.51.0]=[0.0520350.0000000.0260180.0520350.0000000.0000000.0000000.000000]\frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T = \begin{bmatrix} 0.052035 \\ 0.000000 \end{bmatrix} \begin{bmatrix} 1.0 & 0.0 & 0.5 & 1.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.052035} & \mathbf{0.000000} & \mathbf{0.026018} & \mathbf{0.052035} \\ \mathbf{0.000000} & \mathbf{0.000000} & \mathbf{0.000000} & \mathbf{0.000000} \end{bmatrix} ∂L∂b(1)=δ(1)=[0.0520350.000000]\frac{\partial L}{\partial b^{(1)}} = \delta^{(1)} = \begin{bmatrix} \mathbf{0.052035} \\ \mathbf{0.000000} \end{bmatrix}

Phase 4: Single Parameter Update Step (η=0.50\eta = 0.50)

We update all parameters along the direction of steepest descent (−∇L-\nabla L):

1. Layer 2 Updates

Wnew(2)=Wold(2)−η∂L∂W(2)=[1.51.0]−0.5[0.1387600.000000]=[1.4306201.000000]W^{(2)}_{\text{new}} = W^{(2)}_{\text{old}} - \eta \frac{\partial L}{\partial W^{(2)}} = \begin{bmatrix} 1.5 & 1.0 \end{bmatrix} - 0.5 \begin{bmatrix} 0.138760 & 0.000000 \end{bmatrix} = \begin{bmatrix} \mathbf{1.430620} & \mathbf{1.000000} \end{bmatrix} bnew(2)=bold(2)−η∂L∂b(2)=−2.0−0.5(0.034690)=−2.017345b^{(2)}_{\text{new}} = b^{(2)}_{\text{old}} - \eta \frac{\partial L}{\partial b^{(2)}} = -2.0 - 0.5(0.034690) = \mathbf{-2.017345}

2. Layer 1 Updates

Wnew(1)=[3.0−2.00.03.0−2.03.02.0−1.0]−0.5[0.0520350.0000000.0260180.0520350.0000000.0000000.0000000.000000]W^{(1)}_{\text{new}} = \begin{bmatrix} 3.0 & -2.0 & 0.0 & 3.0 \\ -2.0 & 3.0 & 2.0 & -1.0 \end{bmatrix} - 0.5 \begin{bmatrix} 0.052035 & 0.000000 & 0.026018 & 0.052035 \\ 0.000000 & 0.000000 & 0.000000 & 0.000000 \end{bmatrix} Wnew(1)=[2.973982−2.000000−0.0130092.973982−2.0000003.0000002.000000−1.000000]W^{(1)}_{\text{new}} = \begin{bmatrix} \mathbf{2.973982} & \mathbf{-2.000000} & \mathbf{-0.013009} & \mathbf{2.973982} \\ \mathbf{-2.000000} & \mathbf{3.000000} & \mathbf{2.000000} & \mathbf{-1.000000} \end{bmatrix} bnew(1)=[−2.0−1.5]−0.5[0.0520350.000000]=[−2.026018−1.500000]b^{(1)}_{\text{new}} = \begin{bmatrix} -2.0 \\ -1.5 \end{bmatrix} - 0.5 \begin{bmatrix} 0.052035 \\ 0.000000 \end{bmatrix} = \begin{bmatrix} \mathbf{-2.026018} \\ \mathbf{-1.500000} \end{bmatrix}

Phase 5: Second Forward Pass (t=1t = 1) & Proof of Loss Reduction

To prove that the closed cycle reduced error, we re-evaluate the forward pass using the updated parameter matrices:

Step 1: Updated Hidden Layer Forward Pass

Compute znew(1)=Wnew(1)x+bnew(1)z^{(1)}_{\text{new}} = W^{(1)}_{\text{new}} x + b^{(1)}_{\text{new}}:

z1,new(1)=(2.973982)(1.0)+(−2.0)(0.0)+(−0.013009)(0.5)+(2.973982)(1.0)−2.026018z_{1, \text{new}}^{(1)} = (2.973982)(1.0) + (-2.0)(0.0) + (-0.013009)(0.5) + (2.973982)(1.0) - 2.026018 z1,new(1)=2.973982+0.0−0.006505+2.973982−2.026018=+3.915441z_{1, \text{new}}^{(1)} = 2.973982 + 0.0 - 0.006505 + 2.973982 - 2.026018 = \mathbf{+3.915441} z2,new(1)=(−2.0)(1.0)+(3.0)(0.0)+(2.0)(0.5)+(−1.0)(1.0)−1.500000=−3.500000z_{2, \text{new}}^{(1)} = (-2.0)(1.0) + (3.0)(0.0) + (2.0)(0.5) + (-1.0)(1.0) - 1.500000 = \mathbf{-3.500000}

Apply ReLU activation:

anew(1)=[max⁡(0,3.915441)max⁡(0,−3.5)]=[3.9154410.000000]a^{(1)}_{\text{new}} = \begin{bmatrix} \max(0, 3.915441) \\ \max(0, -3.5) \end{bmatrix} = \begin{bmatrix} \mathbf{3.915441} \\ \mathbf{0.000000} \end{bmatrix}

Step 2: Updated Output Layer Forward Pass

Compute znew(2)=Wnew(2)anew(1)+bnew(2)z^{(2)}_{\text{new}} = W^{(2)}_{\text{new}} a^{(1)}_{\text{new}} + b^{(2)}_{\text{new}}:

znew(2)=(1.430620)(3.915441)+(1.0)(0.0)−2.017345=5.601510−2.017345=+3.584165z^{(2)}_{\text{new}} = (1.430620)(3.915441) + (1.0)(0.0) - 2.017345 = 5.601510 - 2.017345 = \mathbf{+3.584165}

Pass through Sigmoid:

anew(2)=σ(3.584165)=11+e−3.584165=11+0.027760≈0.972990(97.30%)a^{(2)}_{\text{new}} = \sigma(3.584165) = \frac{1}{1 + e^{-3.584165}} = \frac{1}{1 + 0.027760} \approx \mathbf{0.972990} \quad (\mathbf{97.30\%})

Step 3: Updated Loss Evaluation

Compute the new Mean Squared Error loss:

L2=(anew(2)−y)2=(0.972990−0.0)2≈0.946710L_2 = (a^{(2)}_{\text{new}} - y)^2 = (0.972990 - 0.0)^2 \approx \mathbf{0.946710}

Algebraic Proof of Strict Loss Reduction

ΔL=L2−L1=0.946710−0.964351=−0.017641<0(Strict Loss Reduction)\Delta L = L_2 - L_1 = 0.946710 - 0.964351 = \mathbf{-0.017641} < 0 \quad (\text{Strict Loss Reduction}) Relative Error Reduction=0.0176410.964351×100%=1.83%\text{Relative Error Reduction} = \frac{0.017641}{0.964351} \times 100\% = \mathbf{1.83\%}

The closed learning cycle successfully nudged the predicted probability from 98.20%98.20\% down to 97.30%97.30\%, strictly reducing Mean Squared Error loss by 1.83%1.83\% on paper in a single update step.


Previous
The Forward-Backward Computational Graph