Μετάβαση στο περιεχόμενο

Αυτό το βίντεο και το κείμενό του είναι στα Αγγλικά.

Mean squared error: the hidden bet behind squaring your mistakes

AI Concepts #018

Squaring errors is not just a habit. Learn what mean squared error quietly assumes about your data, when that bet pays off, and when it misleads you.

Αναλυτικά

What mean squared error is

When a model makes predictions, you need one number that says how wrong it is overall. Mean squared error, or MSE, is the most common choice. For each prediction, take the miss (the gap between the prediction and the real value), square it, and average all those squares. Lower is better, and training a model largely means hunting for the settings that make that number as small as possible.

Squaring makes big misses count far more than small ones. A miss of 2 costs 4. A miss of 10 costs 100.

Most people meet MSE as a convention, the obvious first thing to try. The course makes a stronger claim: squaring is not a habit. It is an assumption.

What the video shows

The video explains that squaring and averaging your mistakes quietly assumes the noise in your measurements is bell-shaped. When that is true, MSE is the right choice. When it is not, it can mislead you.

The hidden assumption

Bell-shaped means the familiar bell curve, also called Gaussian noise: most errors are small, bigger ones are rarer, and really large ones practically never happen.

Start from that belief and ask a natural question: which model settings make the data I actually collected least surprising? Work through the math (the course does it in a few lines), throw away the parts that do not depend on the model, and what survives is exactly the sum of squared errors. So choosing MSE and betting on bell-curve noise are one and the same decision. Anyone who uses it is already making that bet, whether or not anyone told them.

The course checks this on its own measurements. It scans 601 candidate slopes with the full bell-curve score and with plain squared error, and both pick the same best slope, 0.293, down to the last point on the grid.

An everyday example

Imagine stepping on a decent bathroom scale five times in a row. The readings wobble a little around your true weight, sometimes above, sometimes below, never wildly off. That is bell-shaped noise, and averaging with squared error is a sound way to settle on one number.

Now imagine that once in a while the cat jumps on the scale. That is not bell-shaped noise. Squared error punishes big misses hardest, so that single reading gets to drag the answer. That failure has its own name, heavy tails, and its own fix, a robust loss, and both are separate concepts in this series.

A bonus: it measures the noise too

The number left over is useful. If you also let the model estimate how noisy your measuring tool is, the best estimate of that noise level (its variance, the average squared spread) is exactly the mean squared error of the fit. So MSE is both the score you minimise and the noise level that makes your measurements most plausible: a ready answer to "how noisy is my sensor?"

Mean or sum?

Averaging the squares instead of adding them up does not change which settings win. It does change how big each training step is. With a plain sum, feeding the model twice as many examples per step makes every step twice as large. That matters when tuning how a model learns, which the course covers in Chapter 3.

Why it matters

Every loss function, the scoring rule a model is trained to minimise, is really a claim about what your errors look like. Choosing MSE means claiming bell-curve noise. That is often reasonable, but it should be a decision you make on purpose, not a default you inherit.

Learn it step by step in Chapter 2 of our free course AI From Scratch: Where a Loss Function Comes From: Likelihood, Not Convention.

Επίσης στο

Περισσότερα από αυτό το κεφάλαιο

Το κείμενο γράφτηκε με τη βοήθεια AI από το κεφάλαιο 2 του δωρεάν μαθήματός μας.