#009Questo video e il suo testo sono in inglese.
Heavy tails: when rare, huge errors break the usual average
Some errors are tiny almost always, then suddenly enormous. Learn what heavy-tailed noise is, how to spot it, and why the usual averages stop working.
Nel dettaglio
What heavy tails are
Every measurement has some error. We usually picture that error as a bell curve: mostly small misses, a few bigger ones, and huge ones practically never. The tails are the far ends of that curve, where the rare, extreme errors live.
Heavy-tailed noise has fat tails. Errors are small most of the time, but every so often one is enormous, and those enormous errors turn up far more often than a bell curve would ever allow.
What the video shows
The video describes errors that are tiny almost always, then suddenly huge. With that kind of noise, the usual average never settles down: after a million samples it is still climbing. The number that keeps climbing is the variance, the average of the squared errors and the standard way to measure how spread out errors are.
An everyday example
The course takes its example from a workshop. A caliper, a precise measuring tool, reads the width of each part to about a tenth of a millimetre. But once or twice a shift, a chip of swarf (a metal shaving left over from cutting) gets stuck under the jaw of the caliper, and that reading is off by several millimetres.
You see the same shape in daily life. Imagine logging your commute for a month: 20 mornings of 27 minutes, plus 2 mornings of three hours when a road was closed. The average comes out at about 41 minutes, a trip you never actually had, while the middle value stays at 27.
How it works
The course compares two kinds of noise at the same scale: ordinary bell-curve noise, and a heavy-tailed kind modelled by the Cauchy distribution, the standard textbook model for this behaviour. It measures the variance on more and more samples:
| Samples | Bell curve | Heavy-tailed (Cauchy) | |---|---|---| | 100 | 0.0164 | 0.26 | | 1,000 | 0.0146 | 59.9 | | 10,000 | 0.0145 | 358 | | 100,000 | 0.0144 | 3,098 | | 1,000,000 | 0.0144 | 32,886 |
The bell-curve spread settles and stays put. The heavy-tailed spread keeps climbing for as long as you keep sampling, because there is nothing for it to settle on: the Cauchy distribution has no variance, and not even an average. Each new extreme reading drags the estimate up again.
The difference is in how fast the tails thin out. A bell curve's tails fall away brutally fast, so extreme errors almost never appear. The Cauchy's tails fall away slowly, so they keep appearing.
How to spot it
- Measure the spread on growing chunks of your data. A spread that settles is a good sign. One that keeps jumping upward points to heavy tails.
- Look at medians, not just averages. The middle value ignores how extreme the extremes are. The course reports medians in its own experiments for exactly this reason: an average of heavy-tailed errors is itself thrown around by the outliers.
Why it matters
Mean squared error, the most common way to score a model, works by making an average of squared errors as small as possible. With heavy-tailed noise that average does not exist, so the model ends up chasing a quantity that is not there, and a single wild reading can pull the whole answer off course. The fix is a different scoring rule, a robust loss, which is its own concept in this series.
Learn it step by step in Chapter 2 of our free course AI From Scratch: Where a Loss Function Comes From: Likelihood, Not Convention.
Anche su
Altri da questo capitolo
#009
#010Loss landscape: training an AI means finding the bottom of a valley
#011Chain rule of probability: how AI scores a sentence word by word
#012Bayes' rule: reasoning backwards from a clue to its most likely cause
#013Maximum likelihood: the answer that makes your data least surprising
#014Negative log-likelihood: turning a fragile product into a calm sum
Testo scritto con assistenza AI dal capitolo 2 del nostro corso gratuito.