Prediction Error & Loss Functions In Practice hero
LaboratoryPrediction Error & Loss Functions

Prediction Error & Loss Functions In Practice

Master raw error calculation, Mean Squared Error formulas, parabolic curvature geometry, and output loss derivatives through manual calculations.

Work through each problem on paper before expanding the solution details.


Part 1: The Underlying Mechanics Drill

Problem 1: Raw Error Calculation

A loan default prediction model evaluates an applicant and outputs a default probability of y^=0.30\hat{y} = 0.30. Twelve months later, the borrower defaults on the loan (y=1.00y = 1.00).

Calculate the raw prediction error e=y−y^e = y - \hat{y}. State whether the sign indicates an overestimate or an underestimate.

Reveal Solution
e=y−y^=1.00−0.30=+0.70e = y - \hat{y} = 1.00 - 0.30 = \mathbf{+0.70}

Sign Interpretation: The raw error is +0.70+0.70. The positive sign indicates an underestimate—the true outcome exceeded the model's low default probability estimate.

(Under the alternative control-systems convention e=y^−ye = \hat{y} - y, e=0.30−1.00=−0.70e = 0.30 - 1.00 = -0.70.)


Problem 2: Single-Sample Mean Squared Error Loss

A fraud detection network evaluates a legitimate credit card transaction (y=0.00y = 0.00) and produces a false alarm prediction of y^=0.80\hat{y} = 0.80.

Calculate the standard unscaled Mean Squared Error loss L=(y−y^)2L = (y - \hat{y})^2.

Reveal Solution
e=y−y^=0.00−0.80=−0.80e = y - \hat{y} = 0.00 - 0.80 = -0.80 L=(y−y^)2=(−0.80)2=0.64(or 1625)L = (y - \hat{y})^2 = (-0.80)^2 = \mathbf{0.64} \quad \left(\text{or } \frac{16}{25}\right)

Takeaway: The negative sign of the overestimate is eliminated by squaring, producing a positive scalar penalty of 0.640.64.


Problem 3: Scaled MSE Loss Computation (12\frac{1}{2} Convention)

A medical diagnostic model evaluates a patient with a confirmed condition (y=1.00y = 1.00) and outputs a probability score of y^=0.60\hat{y} = 0.60.

Calculate the scaled squared error loss Lscaled=12(y−y^)2L_{\text{scaled}} = \frac{1}{2}(y - \hat{y})^2.

Reveal Solution
e=y−y^=1.00−0.60=+0.40e = y - \hat{y} = 1.00 - 0.60 = +0.40 e2=(0.40)2=0.16e^2 = (0.40)^2 = 0.16 Lscaled=12(0.16)=0.08(or 225)L_{\text{scaled}} = \frac{1}{2}(0.16) = \mathbf{0.08} \quad \left(\text{or } \frac{2}{25}\right)

Takeaway: The scaled loss is exactly half of the standard loss (0.16→0.080.16 \to 0.08), preserving the relative penalty while simplifying future derivative calculations.


Problem 4: Output Loss Derivative Evaluation

A computer vision model inspects a defective solar cell (y=0.00y = 0.00) and outputs a defect-free confidence of y^=0.75\hat{y} = 0.75.

Calculate the exact derivative of standard squared loss with respect to the prediction: ∂L∂y^=2(y^−y)\frac{\partial L}{\partial \hat{y}} = 2(\hat{y} - y). State whether the slope is uphill or downhill, and what action is required to reduce loss.

Reveal Solution
∂L∂y^=2(y^−y)=2(0.75−0.00)=2(0.75)=+1.50\frac{\partial L}{\partial \hat{y}} = 2(\hat{y} - y) = 2(0.75 - 0.00) = 2(0.75) = \mathbf{+1.50}

Interpretation:

  • The derivative is +1.50+1.50 (Positive).
  • A positive slope indicates the model sits on the uphill right wall of the loss bowl: increasing y^\hat{y} further will increase the loss penalty.
  • To reduce loss, optimization must move in the negative gradient direction (−∂L∂y^=−1.50-\frac{\partial L}{\partial \hat{y}} = -1.50), nudging y^\hat{y} downward toward zero.

Problem 5: Parabolic Loss Minimum Stationary Condition

For a target label y=1.00y = 1.00, determine the exact predicted value y^\hat{y} that minimizes the loss function L=(y−y^)2L = (y - \hat{y})^2 by setting the loss derivative to zero: ∂L∂y^=0\frac{\partial L}{\partial \hat{y}} = 0.

Verify the loss value at this optimal point.

Reveal Solution

Set the derivative equal to zero:

∂L∂y^=2(y^−y)=2(y^−1.00)=0\frac{\partial L}{\partial \hat{y}} = 2(\hat{y} - y) = 2(\hat{y} - 1.00) = 0

Divide by 22:

y^−1.00=0  ⟹  y^=1.00\hat{y} - 1.00 = 0 \implies \mathbf{\hat{y} = 1.00}

Evaluate the loss at y^=1.00\hat{y} = 1.00:

L(1.00,1.00)=(1.00−1.00)2=02=0.00L(1.00, 1.00) = (1.00 - 1.00)^2 = 0^2 = \mathbf{0.00}

Geometric Conclusion: The global minimum of the parabolic loss surface occurs precisely where the prediction matches reality (y^=1.00\hat{y} = 1.00). At this point, the tangent slope is zero (∂L∂y^=0.00\frac{\partial L}{\partial \hat{y}} = 0.00), and the penalty reaches its absolute lower bound (L=0.00L = 0.00).


Part 2: Applied Scenario: The VC Pitch Flop Penalty

In Course 1, our venture capital firm uses a neural network to evaluate early-stage startup pitch decks across four locked criteria [Team Experience, Market Size, Competition, Risk].

The investment committee reviewed a hyped fintech venture (OmniCloud Analytics). The network's forward pass produced a highly confident investment greenlight score:

y^=0.85(85% Predicted Probability of Series A Success)\hat{y} = 0.85 \quad (\text{85\% Predicted Probability of Series A Success})

Based on this score, the firm invested $3M\char36 3\text{M}. Eighteen months later, due to severe market saturation and high customer churn, the startup ceased operations, liquidated its assets, and declared bankruptcy:

y=0.00(Observed Ground-Truth Outcome = Bankruptcy Flop)y = 0.00 \quad (\text{Observed Ground-Truth Outcome = Bankruptcy Flop})

Problem 6: Step-by-Step VC Loss & Derivative Audit

  • Part A: Calculate the raw prediction error e=y−y^e = y - \hat{y}.
  • Part B: Calculate the standard Mean Squared Error loss L=(y−y^)2L = (y - \hat{y})^2 and the scaled loss Lscaled=12(y−y^)2L_{\text{scaled}} = \frac{1}{2}(y - \hat{y})^2.
  • Part C: Calculate the output loss derivative under both the standard definition (∂L∂y^=2(y^−y)\frac{\partial L}{\partial \hat{y}} = 2(\hat{y} - y)) and the scaled definition (∂Lscaled∂y^=y^−y\frac{\partial L_{\text{scaled}}}{\partial \hat{y}} = \hat{y} - y).
  • Part D (Investment Post-Mortem): In 2–3 sentences, interpret what the sign and magnitude of the derivative (+1.70+1.70) signal to the upstream neural network weights for future pitch evaluations.
Reveal Solution

Part A: Raw Prediction Error

e=y−y^=0.00−0.85=−0.85e = y - \hat{y} = 0.00 - 0.85 = \mathbf{-0.85}

The negative sign confirms an aggressive over-prediction of +0.85+0.85 above ground-truth reality.


Part B: Standard & Scaled Loss Penalties

  • Standard MSE Loss: L=(y−y^)2=(−0.85)2=0.7225L = (y - \hat{y})^2 = (-0.85)^2 = \mathbf{0.7225}

  • Scaled MSE Loss: Lscaled=12(y−y^)2=12(0.7225)=0.36125L_{\text{scaled}} = \frac{1}{2}(y - \hat{y})^2 = \frac{1}{2}(0.7225) = \mathbf{0.36125}


Part C: Output Loss Derivatives

  • Standard Derivative: ∂L∂y^=2(y^−y)=2(0.85−0.00)=2(0.85)=+1.70\frac{\partial L}{\partial \hat{y}} = 2(\hat{y} - y) = 2(0.85 - 0.00) = 2(0.85) = \mathbf{+1.70}

  • Scaled Derivative: ∂Lscaled∂y^=y^−y=0.85−0.00=+0.85\frac{\partial L_{\text{scaled}}}{\partial \hat{y}} = \hat{y} - y = 0.85 - 0.00 = \mathbf{+0.85}


Part D: Investment Post-Mortem & Parameter Signal

The positive derivative (∂L∂y^=+1.70>0\frac{\partial L}{\partial \hat{y}} = +1.70 > 0) reveals that the model sits high on the right wall of the parabolic loss bowl, where increasing prediction confidence severely escalates financial penalty.

To reduce portfolio loss on similar failing ventures, the negative gradient signal (−∂L∂y^=−1.70-\frac{\partial L}{\partial \hat{y}} = -1.70) will flow backward into the network during backpropagation. This forces the upstream weight matrices to dampen the positive influence of superficial team metrics and heighten sensitivity to crowded competition, pulling future greenlight probabilities downward toward zero.

Previous
Output Loss Derivatives and Gradients