The Mean Squared Error Loss Function hero
Lesson 2Prediction Error & Loss Functions

The Mean Squared Error Loss Function

Formulate the Mean Squared Error loss function to convert prediction mistakes into positive, differentiable scalar penalties for neural networks.

To satisfy all four loss function requirements, we construct the Mean Squared Error (MSE) loss function.


Single-Sample Squared Error Formulation

For a single training example with ground truth yy and prediction y^\hat{y}, the standard squared error loss is defined as:

L(y,y^)=(y−y^)2L(y, \hat{y}) = (y - \hat{y})^2

Because the square of any real number is non-negative, squaring immediately eliminates the cancellation fallacy:

(y−y^)2=(−(y^−y))2=(y^−y)2≥0(y - \hat{y})^2 = \big(-( \hat{y} - y )\big)^2 = (\hat{y} - y)^2 \ge 0

Whether the model overshoots (e=−0.95  ⟹  (−0.95)2=0.9025e = -0.95 \implies (-0.95)^2 = \mathbf{0.9025}) or undershoots (e=+0.95  ⟹  (+0.95)2=0.9025e = +0.95 \implies (+0.95)^2 = \mathbf{0.9025}), both mistakes generate the exact same positive scalar penalty.

Let's re-evaluate our two-movie dataset using squared loss:

  • Movie 1: L1=(1.0−0.50)2=(0.50)2=0.25L_1 = (1.0 - 0.50)^2 = (0.50)^2 = \mathbf{0.25}
  • Movie 2: L2=(0.0−0.50)2=(−0.50)2=0.25L_2 = (0.0 - 0.50)^2 = (-0.50)^2 = \mathbf{0.25}

Summing the squared errors produces a true total penalty:

Ltotal=L1+L2=0.25+0.25=0.50L_{\text{total}} = L_1 + L_2 = 0.25 + 0.25 = \mathbf{0.50}

The errors reinforce rather than cancel, providing an accurate, honest measurement of model inaccuracy.


Quadratic Penalty Scaling: Blunders vs. Small Errors

Why square the error rather than taking the absolute value (∣y−y^∣|y - \hat{y}|, known as Mean Absolute Error or L1 loss)?

Squaring introduces super-linear (quadratic) penalty growth. As the raw error grows linearly, the squared penalty explodes quadratically:

L(e2)L(e1)=(e2e1)2\frac{L(e_2)}{L(e_1)} = \left(\frac{e_2}{e_1}\right)^2

Observe how squared penalty scales across different error magnitudes:

| Error Type | Raw Discrepancy (∣e∣|e|) | Absolute Penalty (∣e∣|e|) | Squared Penalty (e2e^2) | Relative Penalty Growth | | :--- | :--- | :--- | :--- | :--- | | Minor Noise | 0.100.10 | 0.100.10 | 0.010.01 | 1×1\times baseline | | Moderate Error | 0.500.50 (5×5\times) | 0.500.50 | 0.250.25 | 25×25\times penalty | | Major Blunder | 1.001.00 (10×10\times) | 1.001.00 | 1.001.00 | 100×100\times penalty |

A mistake that is 10×10\times larger (1.001.00 vs 0.100.10) produces a squared penalty that is 100×100\times more severe (1.001.00 vs 0.010.01).

Raw Error:      0.10 ────────► 1.00  ( 10x increase )
Squared Loss:   0.01 ────────────────────────► 1.00  ( 100x increase )

This quadratic scaling gives Mean Squared Error a crucial structural property:

  1. Tolerance for Minor Imperfections: Tiny errors (e=0.10→L=0.01e = 0.10 \to L = 0.01) incur negligible penalties, preventing the network from over-reacting to minor noise in training data.
  2. Aggressive Suppression of Blunders: Catastrophic mistakes (e=0.95→L=0.9025e = 0.95 \to L = 0.9025) incur massive penalties, forcing optimization algorithms to focus primary corrective effort on eliminating large blunders first.

TEASER: Beyond Regression: Cross-Entropy for Multi-Class Output

Mean Squared Error is the foundation of regression and continuous prediction. For classification tasks with hundreds of discrete categories (such as language models predicting the next token), models use Cross-Entropy Loss. We will develop Cross-Entropy from first principles in Course 2.


Dataset Mean Squared Error

When evaluating an entire dataset or mini-batch containing NN training examples D={(x(1),y(1)),(x(2),y(2)),…,(x(N),y(N))}\mathcal{D} = \{(x^{(1)}, y^{(1)}), (x^{(2)}, y^{(2)}), \dots, (x^{(N)}, y^{(N)})\}, we take the arithmetic mean of all individual sample losses:

LMSE=1N∑i=1N(y(i)−y^(i))2L_{\text{MSE}} = \frac{1}{N} \sum_{i=1}^{N} \left( y^{(i)} - \hat{y}^{(i)} \right)^2

Dividing by the sample count NN ensures that the loss metric is scale-invariant: a batch of 1,000 samples will produce the same average loss as a batch of 10 samples if the model exhibits the same average prediction accuracy on both.


The 12\frac{1}{2} Scaling Convention in Calculus

When reading machine learning papers and textbooks, you will frequently encounter squared loss written with a constant factor of 12\frac{1}{2}:

Lscaled=12(y−y^)2=12(y^−y)2L_{\text{scaled}} = \frac{1}{2}(y - \hat{y})^2 = \frac{1}{2}(\hat{y} - y)^2

Why include the fraction 12\frac{1}{2}?

In calculus, when we take the derivative of a squared term u2u^2 using the power rule, the exponent 22 drops down as a multiplier:

ddu[u2]=2u\frac{d}{du}[u^2] = 2u

By pre-multiplying the loss function by 12\frac{1}{2}, the 22 from the power rule is canceled out completely:

ddu[12u2]=12⋅2u=u\frac{d}{du}\left[ \frac{1}{2} u^2 \right] = \frac{1}{2} \cdot 2u = \mathbf{u}

This algebraic convenience simplifies subsequent gradient equations, eliminating lingering factors of 22 throughout backpropagation derivations.

Does Scaling by 12\frac{1}{2} Change the Optimal Parameters?

No. Multiplying a loss function by any positive constant c>0c > 0 scales the vertical height of the loss curve, but does not alter the location of its minimum:

arg⁡min⁡θ[12L(θ)]=arg⁡min⁡θL(θ)\arg\min_{\theta} \left[ \frac{1}{2} L(\theta) \right] = \arg\min_{\theta} L(\theta)

The parameter values that minimize 12L\frac{1}{2}L are identical to the parameter values that minimize LL. In gradient descent, any constant scalar factor is absorbed directly into the learning rate hyperparameter (η\eta).

NOTE: Reporting vs. Training Loss

In production machine learning frameworks:

  • Standard MSE (L=1N∑(y−y^)2L = \frac{1}{N}\sum(y-\hat{y})^2): Reported in evaluation benchmarks and test metrics because its square root (RMSE=MSE\text{RMSE} = \sqrt{\text{MSE}}) directly reflects error in the original physical units of the target variable.
  • Scaled MSE (L=12N∑(y−y^)2L = \frac{1}{2N}\sum(y-\hat{y})^2): Frequently used internally inside optimization code to simplify analytical gradients.

Previous
Grounding Prediction Error in Reality