Let us trace the complete, unbroken learning cycle on paper with explicit numbers for our running 2-layer network.
System Configuration: 'Die Hard in Space'
We evaluate the screenplay for 'Die Hard in Space' across our running 4 → 2 → 1 4 \to 2 \to 1 4 → 2 → 1 Multi-Layer Perceptron.
Input Feature Vector (x ∈ R 4 × 1 x \in \mathbb{R}^{4 \times 1} x ∈ R 4 × 1 ):
x = [ 1.0 0.0 0.5 1.0 ] (Action Measurement) (Romance Measurement) (Comedy Measurement) (Sci-Fi Measurement) x = \begin{bmatrix} 1.0 \\ 0.0 \\ 0.5 \\ 1.0 \end{bmatrix} \begin{matrix} \text{(Action Measurement)} \\ \text{(Romance Measurement)} \\ \text{(Comedy Measurement)} \\ \text{(Sci-Fi Measurement)} \end{matrix} x = 1.0 0.0 0.5 1.0 (Action Measurement) (Romance Measurement) (Comedy Measurement) (Sci-Fi Measurement)
Layer 1 Initial Parameters (W ( 1 ) ∈ R 2 × 4 , b ( 1 ) ∈ R 2 × 1 W^{(1)} \in \mathbb{R}^{2 \times 4}, b^{(1)} \in \mathbb{R}^{2 \times 1} W ( 1 ) ∈ R 2 × 4 , b ( 1 ) ∈ R 2 × 1 ):
W ( 1 ) = [ 3.0 − 2.0 0.0 3.0 − 2.0 3.0 2.0 − 1.0 ] , b ( 1 ) = [ − 2.0 − 1.5 ] W^{(1)} = \begin{bmatrix} 3.0 & -2.0 & 0.0 & 3.0 \\ -2.0 & 3.0 & 2.0 & -1.0 \end{bmatrix}, \qquad b^{(1)} = \begin{bmatrix} -2.0 \\ -1.5 \end{bmatrix} W ( 1 ) = [ 3.0 − 2.0 − 2.0 3.0 0.0 2.0 3.0 − 1.0 ] , b ( 1 ) = [ − 2.0 − 1.5 ]
Hidden activation function: f ( z ) = ReLU ( z ) = max ( 0 , z ) f(z) = \text{ReLU}(z) = \max(0, z) f ( z ) = ReLU ( z ) = max ( 0 , z ) .
Layer 2 Initial Parameters (W ( 2 ) ∈ R 1 × 2 , b ( 2 ) ∈ R 1 × 1 W^{(2)} \in \mathbb{R}^{1 \times 2}, b^{(2)} \in \mathbb{R}^{1 \times 1} W ( 2 ) ∈ R 1 × 2 , b ( 2 ) ∈ R 1 × 1 ):
W ( 2 ) = [ 1.5 1.0 ] , b ( 2 ) = [ − 2.0 ] W^{(2)} = \begin{bmatrix} 1.5 & 1.0 \end{bmatrix}, \qquad b^{(2)} = \begin{bmatrix} -2.0 \end{bmatrix} W ( 2 ) = [ 1.5 1.0 ] , b ( 2 ) = [ − 2.0 ]
Output activation function: σ ( z ) = 1 1 + e − z \sigma(z) = \frac{1}{1 + e^{-z}} σ ( z ) = 1 + e − z 1 .
Ground-Truth Target: y = 0.0 y = 0.0 y = 0.0 (Box Office Flop).
Step Size Factor: η = 0.50 \eta = 0.50 η = 0.50 .
Phase 1: Initial Forward Pass (t = 0 t = 0 t = 0 )
Step 1: Layer 1 Affine Transformation
Compute z ( 1 ) = W ( 1 ) x + b ( 1 ) z^{(1)} = W^{(1)} x + b^{(1)} z ( 1 ) = W ( 1 ) x + b ( 1 ) :
z 1 ( 1 ) = ( 3.0 ) ( 1.0 ) + ( − 2.0 ) ( 0.0 ) + ( 0.0 ) ( 0.5 ) + ( 3.0 ) ( 1.0 ) − 2.0 = 3.0 + 0.0 + 0.0 + 3.0 − 2.0 = + 4.0 z_1^{(1)} = (3.0)(1.0) + (-2.0)(0.0) + (0.0)(0.5) + (3.0)(1.0) - 2.0 = 3.0 + 0.0 + 0.0 + 3.0 - 2.0 = \mathbf{+4.0} z 1 ( 1 ) = ( 3.0 ) ( 1.0 ) + ( − 2.0 ) ( 0.0 ) + ( 0.0 ) ( 0.5 ) + ( 3.0 ) ( 1.0 ) − 2.0 = 3.0 + 0.0 + 0.0 + 3.0 − 2.0 = + 4.0
z 2 ( 1 ) = ( − 2.0 ) ( 1.0 ) + ( 3.0 ) ( 0.0 ) + ( 2.0 ) ( 0.5 ) + ( − 1.0 ) ( 1.0 ) − 1.5 = − 2.0 + 0.0 + 1.0 − 1.0 − 1.5 = − 3.5 z_2^{(1)} = (-2.0)(1.0) + (3.0)(0.0) + (2.0)(0.5) + (-1.0)(1.0) - 1.5 = -2.0 + 0.0 + 1.0 - 1.0 - 1.5 = \mathbf{-3.5} z 2 ( 1 ) = ( − 2.0 ) ( 1.0 ) + ( 3.0 ) ( 0.0 ) + ( 2.0 ) ( 0.5 ) + ( − 1.0 ) ( 1.0 ) − 1.5 = − 2.0 + 0.0 + 1.0 − 1.0 − 1.5 = − 3.5
z ( 1 ) = [ + 4.0 − 3.5 ] z^{(1)} = \begin{bmatrix} +4.0 \\ -3.5 \end{bmatrix} z ( 1 ) = [ + 4.0 − 3.5 ]
Step 2: Layer 1 Non-Linear Activation
Apply element-wise ReLU activation:
a ( 1 ) = ReLU ( z ( 1 ) ) = [ max ( 0 , 4.0 ) max ( 0 , − 3.5 ) ] = [ 4.0 0.0 ] (Summer Popcorn Flick Factor) (Rom-Com Factor) a^{(1)} = \text{ReLU}\left(z^{(1)}\right) = \begin{bmatrix} \max(0, 4.0) \\ \max(0, -3.5) \end{bmatrix} = \begin{bmatrix} \mathbf{4.0} \\ \mathbf{0.0} \end{bmatrix} \begin{matrix} \text{(Summer Popcorn Flick Factor)} \\ \text{(Rom-Com Factor)} \end{matrix} a ( 1 ) = ReLU ( z ( 1 ) ) = [ max ( 0 , 4.0 ) max ( 0 , − 3.5 ) ] = [ 4.0 0.0 ] (Summer Popcorn Flick Factor) (Rom-Com Factor)
Step 3: Layer 2 Affine Transformation
Compute scalar pre-activation z ( 2 ) = W ( 2 ) a ( 1 ) + b ( 2 ) z^{(2)} = W^{(2)} a^{(1)} + b^{(2)} z ( 2 ) = W ( 2 ) a ( 1 ) + b ( 2 ) :
z ( 2 ) = ( 1.5 ) ( 4.0 ) + ( 1.0 ) ( 0.0 ) + ( − 2.0 ) = 6.0 + 0.0 − 2.0 = + 4.0 z^{(2)} = (1.5)(4.0) + (1.0)(0.0) + (-2.0) = 6.0 + 0.0 - 2.0 = \mathbf{+4.0} z ( 2 ) = ( 1.5 ) ( 4.0 ) + ( 1.0 ) ( 0.0 ) + ( − 2.0 ) = 6.0 + 0.0 − 2.0 = + 4.0
Step 4: Layer 2 Sigmoid Probability Squashing
Compute final prediction y ^ = a ( 2 ) = σ ( z ( 2 ) ) \hat{y} = a^{(2)} = \sigma(z^{(2)}) y ^ = a ( 2 ) = σ ( z ( 2 ) ) :
a ( 2 ) = σ ( 4.0 ) = 1 1 + e − 4.0 = 1 1 + 0.018316 ≈ 0.982014 ( 98.20 % ) a^{(2)} = \sigma(4.0) = \frac{1}{1 + e^{-4.0}} = \frac{1}{1 + 0.018316} \approx \mathbf{0.982014} \quad (\mathbf{98.20\%}) a ( 2 ) = σ ( 4.0 ) = 1 + e − 4.0 1 = 1 + 0.018316 1 ≈ 0.982014 ( 98.20% )
Phase 2: Error Quantification & Loss Evaluation
The network predicted a 98.20 % 98.20\% 98.20% Hit probability, but the movie was a box office Flop (y = 0.0 y = 0.0 y = 0.0 ).
Step 1: Raw Prediction Error
e = y − y ^ = 0.0 − 0.982014 = − 0.982014 e = y - \hat{y} = 0.0 - 0.982014 = \mathbf{-0.982014} e = y − y ^ = 0.0 − 0.982014 = − 0.982014
Step 2: Mean Squared Error Loss
L 1 = ( a ( 2 ) − y ) 2 = ( 0.982014 − 0.0 ) 2 ≈ 0.964351 L_1 = (a^{(2)} - y)^2 = (0.982014 - 0.0)^2 \approx \mathbf{0.964351} L 1 = ( a ( 2 ) − y ) 2 = ( 0.982014 − 0.0 ) 2 ≈ 0.964351
Phase 3: Backward Pass & Gradient Attribution
Step 1: Output Layer Error Attribution
Loss Sensitivity:
∂ L ∂ a ( 2 ) = 2 ( a ( 2 ) − y ) = 2 ( 0.982014 − 0.0 ) = + 1.964028 \frac{\partial L}{\partial a^{(2)}} = 2(a^{(2)} - y) = 2(0.982014 - 0.0) = \mathbf{+1.964028} ∂ a ( 2 ) ∂ L = 2 ( a ( 2 ) − y ) = 2 ( 0.982014 − 0.0 ) = + 1.964028
Activation Function Slope:
σ ′ ( z ( 2 ) ) = a ( 2 ) ( 1 − a ( 2 ) ) = 0.982014 × ( 1 − 0.982014 ) = 0.982014 × 0.017986 = 0.017663 \sigma'(z^{(2)}) = a^{(2)}(1 - a^{(2)}) = 0.982014 \times (1 - 0.982014) = 0.982014 \times 0.017986 = \mathbf{0.017663} σ ′ ( z ( 2 ) ) = a ( 2 ) ( 1 − a ( 2 ) ) = 0.982014 × ( 1 − 0.982014 ) = 0.982014 × 0.017986 = 0.017663
Output Error Delta (δ ( 2 ) \delta^{(2)} δ ( 2 ) ):
δ ( 2 ) = ∂ L ∂ a ( 2 ) ⋅ σ ′ ( z ( 2 ) ) = ( + 1.964028 ) × ( 0.017663 ) = + 0.034690 \delta^{(2)} = \frac{\partial L}{\partial a^{(2)}} \cdot \sigma'(z^{(2)}) = (+1.964028) \times (0.017663) = \mathbf{+0.034690} δ ( 2 ) = ∂ a ( 2 ) ∂ L ⋅ σ ′ ( z ( 2 ) ) = ( + 1.964028 ) × ( 0.017663 ) = + 0.034690
Output Parameter Gradients:
∂ L ∂ W ( 2 ) = δ ( 2 ) ( a ( 1 ) ) T = 0.034690 × [ 4.0 0.0 ] = [ 0.138760 0.000000 ] \frac{\partial L}{\partial W^{(2)}} = \delta^{(2)} (a^{(1)})^T = 0.034690 \times \begin{bmatrix} 4.0 & 0.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.138760} & \mathbf{0.000000} \end{bmatrix} ∂ W ( 2 ) ∂ L = δ ( 2 ) ( a ( 1 ) ) T = 0.034690 × [ 4.0 0.0 ] = [ 0.138760 0.000000 ]
∂ L ∂ b ( 2 ) = δ ( 2 ) = + 0.034690 \frac{\partial L}{\partial b^{(2)}} = \delta^{(2)} = \mathbf{+0.034690} ∂ b ( 2 ) ∂ L = δ ( 2 ) = + 0.034690
Step 2: Hidden Layer Error Backpropagation
Transposed Weight Projection:
( W ( 2 ) ) T δ ( 2 ) = [ 1.5 1.0 ] ( 0.034690 ) = [ 1.5 × 0.034690 1.0 × 0.034690 ] = [ 0.052035 0.034690 ] (W^{(2)})^T \delta^{(2)} = \begin{bmatrix} 1.5 \\ 1.0 \end{bmatrix} (0.034690) = \begin{bmatrix} 1.5 \times 0.034690 \\ 1.0 \times 0.034690 \end{bmatrix} = \begin{bmatrix} \mathbf{0.052035} \\ \mathbf{0.034690} \end{bmatrix} ( W ( 2 ) ) T δ ( 2 ) = [ 1.5 1.0 ] ( 0.034690 ) = [ 1.5 × 0.034690 1.0 × 0.034690 ] = [ 0.052035 0.034690 ]
Hidden Activation Slopes:
f ′ ( z ( 1 ) ) = [ ReLU ′ ( + 4.0 ) ReLU ′ ( − 3.5 ) ] = [ 1.0 0.0 ] f'(z^{(1)}) = \begin{bmatrix} \text{ReLU}'(+4.0) \\ \text{ReLU}'(-3.5) \end{bmatrix} = \begin{bmatrix} \mathbf{1.0} \\ \mathbf{0.0} \end{bmatrix} f ′ ( z ( 1 ) ) = [ ReLU ′ ( + 4.0 ) ReLU ′ ( − 3.5 ) ] = [ 1.0 0.0 ]
Hidden Error Delta (δ ( 1 ) \delta^{(1)} δ ( 1 ) ):
δ ( 1 ) = ( ( W ( 2 ) ) T δ ( 2 ) ) ⊙ f ′ ( z ( 1 ) ) = [ 0.052035 0.034690 ] ⊙ [ 1.0 0.0 ] = [ 0.052035 0.000000 ] \delta^{(1)} = \left( (W^{(2)})^T \delta^{(2)} \right) \odot f'(z^{(1)}) = \begin{bmatrix} 0.052035 \\ 0.034690 \end{bmatrix} \odot \begin{bmatrix} 1.0 \\ 0.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.052035} \\ \mathbf{0.000000} \end{bmatrix} δ ( 1 ) = ( ( W ( 2 ) ) T δ ( 2 ) ) ⊙ f ′ ( z ( 1 ) ) = [ 0.052035 0.034690 ] ⊙ [ 1.0 0.0 ] = [ 0.052035 0.000000 ]
Hidden Layer Parameter Gradients:
∂ L ∂ W ( 1 ) = δ ( 1 ) x T = [ 0.052035 0.000000 ] [ 1.0 0.0 0.5 1.0 ] = [ 0.052035 0.000000 0.026018 0.052035 0.000000 0.000000 0.000000 0.000000 ] \frac{\partial L}{\partial W^{(1)}} = \delta^{(1)} x^T = \begin{bmatrix} 0.052035 \\ 0.000000 \end{bmatrix} \begin{bmatrix} 1.0 & 0.0 & 0.5 & 1.0 \end{bmatrix} = \begin{bmatrix} \mathbf{0.052035} & \mathbf{0.000000} & \mathbf{0.026018} & \mathbf{0.052035} \\ \mathbf{0.000000} & \mathbf{0.000000} & \mathbf{0.000000} & \mathbf{0.000000} \end{bmatrix} ∂ W ( 1 ) ∂ L = δ ( 1 ) x T = [ 0.052035 0.000000 ] [ 1.0 0.0 0.5 1.0 ] = [ 0.052035 0.000000 0.000000 0.000000 0.026018 0.000000 0.052035 0.000000 ]
∂ L ∂ b ( 1 ) = δ ( 1 ) = [ 0.052035 0.000000 ] \frac{\partial L}{\partial b^{(1)}} = \delta^{(1)} = \begin{bmatrix} \mathbf{0.052035} \\ \mathbf{0.000000} \end{bmatrix} ∂ b ( 1 ) ∂ L = δ ( 1 ) = [ 0.052035 0.000000 ]
Phase 4: Single Parameter Update Step (η = 0.50 \eta = 0.50 η = 0.50 )
We update all parameters along the direction of steepest descent (− ∇ L -\nabla L − ∇ L ):
1. Layer 2 Updates
W new ( 2 ) = W old ( 2 ) − η ∂ L ∂ W ( 2 ) = [ 1.5 1.0 ] − 0.5 [ 0.138760 0.000000 ] = [ 1.430620 1.000000 ] W^{(2)}_{\text{new}} = W^{(2)}_{\text{old}} - \eta \frac{\partial L}{\partial W^{(2)}} = \begin{bmatrix} 1.5 & 1.0 \end{bmatrix} - 0.5 \begin{bmatrix} 0.138760 & 0.000000 \end{bmatrix} = \begin{bmatrix} \mathbf{1.430620} & \mathbf{1.000000} \end{bmatrix} W new ( 2 ) = W old ( 2 ) − η ∂ W ( 2 ) ∂ L = [ 1.5 1.0 ] − 0.5 [ 0.138760 0.000000 ] = [ 1.430620 1.000000 ]
b new ( 2 ) = b old ( 2 ) − η ∂ L ∂ b ( 2 ) = − 2.0 − 0.5 ( 0.034690 ) = − 2.017345 b^{(2)}_{\text{new}} = b^{(2)}_{\text{old}} - \eta \frac{\partial L}{\partial b^{(2)}} = -2.0 - 0.5(0.034690) = \mathbf{-2.017345} b new ( 2 ) = b old ( 2 ) − η ∂ b ( 2 ) ∂ L = − 2.0 − 0.5 ( 0.034690 ) = − 2.017345
2. Layer 1 Updates
W new ( 1 ) = [ 3.0 − 2.0 0.0 3.0 − 2.0 3.0 2.0 − 1.0 ] − 0.5 [ 0.052035 0.000000 0.026018 0.052035 0.000000 0.000000 0.000000 0.000000 ] W^{(1)}_{\text{new}} = \begin{bmatrix} 3.0 & -2.0 & 0.0 & 3.0 \\ -2.0 & 3.0 & 2.0 & -1.0 \end{bmatrix} - 0.5 \begin{bmatrix} 0.052035 & 0.000000 & 0.026018 & 0.052035 \\ 0.000000 & 0.000000 & 0.000000 & 0.000000 \end{bmatrix} W new ( 1 ) = [ 3.0 − 2.0 − 2.0 3.0 0.0 2.0 3.0 − 1.0 ] − 0.5 [ 0.052035 0.000000 0.000000 0.000000 0.026018 0.000000 0.052035 0.000000 ]
W new ( 1 ) = [ 2.973982 − 2.000000 − 0.013009 2.973982 − 2.000000 3.000000 2.000000 − 1.000000 ] W^{(1)}_{\text{new}} = \begin{bmatrix} \mathbf{2.973982} & \mathbf{-2.000000} & \mathbf{-0.013009} & \mathbf{2.973982} \\ \mathbf{-2.000000} & \mathbf{3.000000} & \mathbf{2.000000} & \mathbf{-1.000000} \end{bmatrix} W new ( 1 ) = [ 2.973982 − 2.000000 − 2.000000 3.000000 − 0.013009 2.000000 2.973982 − 1.000000 ]
b new ( 1 ) = [ − 2.0 − 1.5 ] − 0.5 [ 0.052035 0.000000 ] = [ − 2.026018 − 1.500000 ] b^{(1)}_{\text{new}} = \begin{bmatrix} -2.0 \\ -1.5 \end{bmatrix} - 0.5 \begin{bmatrix} 0.052035 \\ 0.000000 \end{bmatrix} = \begin{bmatrix} \mathbf{-2.026018} \\ \mathbf{-1.500000} \end{bmatrix} b new ( 1 ) = [ − 2.0 − 1.5 ] − 0.5 [ 0.052035 0.000000 ] = [ − 2.026018 − 1.500000 ]
Phase 5: Second Forward Pass (t = 1 t = 1 t = 1 ) & Proof of Loss Reduction
To prove that the closed cycle reduced error, we re-evaluate the forward pass using the updated parameter matrices:
Step 1: Updated Hidden Layer Forward Pass
Compute z new ( 1 ) = W new ( 1 ) x + b new ( 1 ) z^{(1)}_{\text{new}} = W^{(1)}_{\text{new}} x + b^{(1)}_{\text{new}} z new ( 1 ) = W new ( 1 ) x + b new ( 1 ) :
z 1 , new ( 1 ) = ( 2.973982 ) ( 1.0 ) + ( − 2.0 ) ( 0.0 ) + ( − 0.013009 ) ( 0.5 ) + ( 2.973982 ) ( 1.0 ) − 2.026018 z_{1, \text{new}}^{(1)} = (2.973982)(1.0) + (-2.0)(0.0) + (-0.013009)(0.5) + (2.973982)(1.0) - 2.026018 z 1 , new ( 1 ) = ( 2.973982 ) ( 1.0 ) + ( − 2.0 ) ( 0.0 ) + ( − 0.013009 ) ( 0.5 ) + ( 2.973982 ) ( 1.0 ) − 2.026018
z 1 , new ( 1 ) = 2.973982 + 0.0 − 0.006505 + 2.973982 − 2.026018 = + 3.915441 z_{1, \text{new}}^{(1)} = 2.973982 + 0.0 - 0.006505 + 2.973982 - 2.026018 = \mathbf{+3.915441} z 1 , new ( 1 ) = 2.973982 + 0.0 − 0.006505 + 2.973982 − 2.026018 = + 3.915441
z 2 , new ( 1 ) = ( − 2.0 ) ( 1.0 ) + ( 3.0 ) ( 0.0 ) + ( 2.0 ) ( 0.5 ) + ( − 1.0 ) ( 1.0 ) − 1.500000 = − 3.500000 z_{2, \text{new}}^{(1)} = (-2.0)(1.0) + (3.0)(0.0) + (2.0)(0.5) + (-1.0)(1.0) - 1.500000 = \mathbf{-3.500000} z 2 , new ( 1 ) = ( − 2.0 ) ( 1.0 ) + ( 3.0 ) ( 0.0 ) + ( 2.0 ) ( 0.5 ) + ( − 1.0 ) ( 1.0 ) − 1.500000 = − 3.500000
Apply ReLU activation:
a new ( 1 ) = [ max ( 0 , 3.915441 ) max ( 0 , − 3.5 ) ] = [ 3.915441 0.000000 ] a^{(1)}_{\text{new}} = \begin{bmatrix} \max(0, 3.915441) \\ \max(0, -3.5) \end{bmatrix} = \begin{bmatrix} \mathbf{3.915441} \\ \mathbf{0.000000} \end{bmatrix} a new ( 1 ) = [ max ( 0 , 3.915441 ) max ( 0 , − 3.5 ) ] = [ 3.915441 0.000000 ]
Step 2: Updated Output Layer Forward Pass
Compute z new ( 2 ) = W new ( 2 ) a new ( 1 ) + b new ( 2 ) z^{(2)}_{\text{new}} = W^{(2)}_{\text{new}} a^{(1)}_{\text{new}} + b^{(2)}_{\text{new}} z new ( 2 ) = W new ( 2 ) a new ( 1 ) + b new ( 2 ) :
z new ( 2 ) = ( 1.430620 ) ( 3.915441 ) + ( 1.0 ) ( 0.0 ) − 2.017345 = 5.601510 − 2.017345 = + 3.584165 z^{(2)}_{\text{new}} = (1.430620)(3.915441) + (1.0)(0.0) - 2.017345 = 5.601510 - 2.017345 = \mathbf{+3.584165} z new ( 2 ) = ( 1.430620 ) ( 3.915441 ) + ( 1.0 ) ( 0.0 ) − 2.017345 = 5.601510 − 2.017345 = + 3.584165
Pass through Sigmoid:
a new ( 2 ) = σ ( 3.584165 ) = 1 1 + e − 3.584165 = 1 1 + 0.027760 ≈ 0.972990 ( 97.30 % ) a^{(2)}_{\text{new}} = \sigma(3.584165) = \frac{1}{1 + e^{-3.584165}} = \frac{1}{1 + 0.027760} \approx \mathbf{0.972990} \quad (\mathbf{97.30\%}) a new ( 2 ) = σ ( 3.584165 ) = 1 + e − 3.584165 1 = 1 + 0.027760 1 ≈ 0.972990 ( 97.30% )
Step 3: Updated Loss Evaluation
Compute the new Mean Squared Error loss:
L 2 = ( a new ( 2 ) − y ) 2 = ( 0.972990 − 0.0 ) 2 ≈ 0.946710 L_2 = (a^{(2)}_{\text{new}} - y)^2 = (0.972990 - 0.0)^2 \approx \mathbf{0.946710} L 2 = ( a new ( 2 ) − y ) 2 = ( 0.972990 − 0.0 ) 2 ≈ 0.946710
Algebraic Proof of Strict Loss Reduction
Δ L = L 2 − L 1 = 0.946710 − 0.964351 = − 0.017641 < 0 ( Strict Loss Reduction ) \Delta L = L_2 - L_1 = 0.946710 - 0.964351 = \mathbf{-0.017641} < 0 \quad (\text{Strict Loss Reduction}) Δ L = L 2 − L 1 = 0.946710 − 0.964351 = − 0.017641 < 0 ( Strict Loss Reduction )
Relative Error Reduction = 0.017641 0.964351 × 100 % = 1.83 % \text{Relative Error Reduction} = \frac{0.017641}{0.964351} \times 100\% = \mathbf{1.83\%} Relative Error Reduction = 0.964351 0.017641 × 100% = 1.83%
The closed learning cycle successfully nudged the predicted probability from 98.20 % 98.20\% 98.20% down to 97.30 % 97.30\% 97.30% , strictly reducing Mean Squared Error loss by 1.83 % 1.83\% 1.83% on paper in a single update step.