
The Mean Squared Error Loss Function
Formulate the Mean Squared Error loss function to convert prediction mistakes into positive, differentiable scalar penalties for neural networks.
To satisfy all four loss function requirements, we construct the Mean Squared Error (MSE) loss function.
Single-Sample Squared Error Formulation
For a single training example with ground truth and prediction , the standard squared error loss is defined as:
Because the square of any real number is non-negative, squaring immediately eliminates the cancellation fallacy:
Whether the model overshoots () or undershoots (), both mistakes generate the exact same positive scalar penalty.
Let's re-evaluate our two-movie dataset using squared loss:
- Movie 1:
- Movie 2:
Summing the squared errors produces a true total penalty:
The errors reinforce rather than cancel, providing an accurate, honest measurement of model inaccuracy.
Quadratic Penalty Scaling: Blunders vs. Small Errors
Why square the error rather than taking the absolute value (, known as Mean Absolute Error or L1 loss)?
Squaring introduces super-linear (quadratic) penalty growth. As the raw error grows linearly, the squared penalty explodes quadratically:
Observe how squared penalty scales across different error magnitudes:
| Error Type | Raw Discrepancy () | Absolute Penalty () | Squared Penalty () | Relative Penalty Growth | | :--- | :--- | :--- | :--- | :--- | | Minor Noise | | | | baseline | | Moderate Error | () | | | penalty | | Major Blunder | () | | | penalty |
A mistake that is larger ( vs ) produces a squared penalty that is more severe ( vs ).
Raw Error: 0.10 ────────► 1.00 ( 10x increase )
Squared Loss: 0.01 ────────────────────────► 1.00 ( 100x increase )
This quadratic scaling gives Mean Squared Error a crucial structural property:
- Tolerance for Minor Imperfections: Tiny errors () incur negligible penalties, preventing the network from over-reacting to minor noise in training data.
- Aggressive Suppression of Blunders: Catastrophic mistakes () incur massive penalties, forcing optimization algorithms to focus primary corrective effort on eliminating large blunders first.
TEASER: Beyond Regression: Cross-Entropy for Multi-Class Output
Mean Squared Error is the foundation of regression and continuous prediction. For classification tasks with hundreds of discrete categories (such as language models predicting the next token), models use Cross-Entropy Loss. We will develop Cross-Entropy from first principles in Course 2.
Dataset Mean Squared Error
When evaluating an entire dataset or mini-batch containing training examples , we take the arithmetic mean of all individual sample losses:
Dividing by the sample count ensures that the loss metric is scale-invariant: a batch of 1,000 samples will produce the same average loss as a batch of 10 samples if the model exhibits the same average prediction accuracy on both.
The Scaling Convention in Calculus
When reading machine learning papers and textbooks, you will frequently encounter squared loss written with a constant factor of :
Why include the fraction ?
In calculus, when we take the derivative of a squared term using the power rule, the exponent drops down as a multiplier:
By pre-multiplying the loss function by , the from the power rule is canceled out completely:
This algebraic convenience simplifies subsequent gradient equations, eliminating lingering factors of throughout backpropagation derivations.
Does Scaling by Change the Optimal Parameters?
No. Multiplying a loss function by any positive constant scales the vertical height of the loss curve, but does not alter the location of its minimum:
The parameter values that minimize are identical to the parameter values that minimize . In gradient descent, any constant scalar factor is absorbed directly into the learning rate hyperparameter ().
NOTE: Reporting vs. Training Loss
In production machine learning frameworks:
- Standard MSE (): Reported in evaluation benchmarks and test metrics because its square root () directly reflects error in the original physical units of the target variable.
- Scaled MSE (): Frequently used internally inside optimization code to simplify analytical gradients.