Saltar ao contido

Este vídeo e o seu texto están en inglés.

Heavy tails: when rare, huge errors break the usual average

AI Concepts #019

Some errors are tiny almost always, then suddenly enormous. Learn what heavy-tailed noise is, how to spot it, and why the usual averages stop working.

En detalle

What heavy tails are

Every measurement has some error. We usually picture that error as a bell curve: mostly small misses, a few bigger ones, and huge ones practically never. The tails are the far ends of that curve, where the rare, extreme errors live.

Heavy-tailed noise has fat tails. Errors are small most of the time, but every so often one is enormous, and those enormous errors turn up far more often than a bell curve would ever allow.

What the video shows

The video describes errors that are tiny almost always, then suddenly huge. With that kind of noise, the usual average never settles down: after a million samples it is still climbing. The number that keeps climbing is the variance, the average of the squared errors and the standard way to measure how spread out errors are.

An everyday example

The course takes its example from a workshop. A caliper, a precise measuring tool, reads the width of each part to about a tenth of a millimetre. But once or twice a shift, a chip of swarf (a metal shaving left over from cutting) gets stuck under the jaw of the caliper, and that reading is off by several millimetres.

You see the same shape in daily life. Imagine logging your commute for a month: 20 mornings of 27 minutes, plus 2 mornings of three hours when a road was closed. The average comes out at about 41 minutes, a trip you never actually had, while the middle value stays at 27.

How it works

The course compares two kinds of noise at the same scale: ordinary bell-curve noise, and a heavy-tailed kind modelled by the Cauchy distribution, the standard textbook model for this behaviour. It measures the variance on more and more samples:

| Samples | Bell curve | Heavy-tailed (Cauchy) | |---|---|---| | 100 | 0.0164 | 0.26 | | 1,000 | 0.0146 | 59.9 | | 10,000 | 0.0145 | 358 | | 100,000 | 0.0144 | 3,098 | | 1,000,000 | 0.0144 | 32,886 |

The bell-curve spread settles and stays put. The heavy-tailed spread keeps climbing for as long as you keep sampling, because there is nothing for it to settle on: the Cauchy distribution has no variance, and not even an average. Each new extreme reading drags the estimate up again.

The difference is in how fast the tails thin out. A bell curve's tails fall away brutally fast, so extreme errors almost never appear. The Cauchy's tails fall away slowly, so they keep appearing.

How to spot it

  • Measure the spread on growing chunks of your data. A spread that settles is a good sign. One that keeps jumping upward points to heavy tails.
  • Look at medians, not just averages. The middle value ignores how extreme the extremes are. The course reports medians in its own experiments for exactly this reason: an average of heavy-tailed errors is itself thrown around by the outliers.

Why it matters

Mean squared error, the most common way to score a model, works by making an average of squared errors as small as possible. With heavy-tailed noise that average does not exist, so the model ends up chasing a quantity that is not there, and a single wild reading can pull the whole answer off course. The fix is a different scoring rule, a robust loss, which is its own concept in this series.

Learn it step by step in Chapter 2 of our free course AI From Scratch: Where a Loss Function Comes From: Likelihood, Not Convention.

Tamén en

Máis deste capítulo

Texto escrito coa axuda da IA a partir do capítulo 2 do noso curso gratuíto.