#01Cette vidéo et son texte sont en anglais.
Input normalization: three lines of code, 6,000 times faster training
Subtracting the average from your data let the same simple model train in 2 passes instead of 11,976. Learn why moving the zero point helps so much.
En détail
What input normalization is
Input normalization means preparing your numbers before a model learns from them, so they sit in a tidy, sensible range. The version in this video is the simplest one, called centring: work out the average of each measurement, then subtract that average from every example. Afterwards, the average sits exactly at zero, with some examples a little above it and some a little below. (Many projects also rescale each measurement so they all spread out by a similar amount; the course's example only needs the first step.)
Nothing about the data is lost. Every difference between two examples stays exactly the same. Only the reference point moves.
What the video shows
The video's hook: three lines of code made the same simple model learn about 6,000 times faster. "Faster" is counted in passes through the data (one pass means looking at every example once). Shift the data so it sits around zero before training, and the passes needed drop from 11,976 to just 2.
An everyday example
Imagine a school built on a hill 1,000 metres above sea level that records each student's height as "height above sea level". Every entry looks like 1,001.60 m or 1,001.74 m. The differences you actually care about are buried in the last digits. Now subtract the class average from each entry and you get small numbers like −7 cm or +7 cm. Same information, but the differences are now the whole number instead of a detail at the end of it.
How it works
In the course, a perceptron (one of the simplest learning programs, which splits examples into two groups with one straight line) sorts eight factory parts using width and weight. Raw, the parts form a small cloud floating far from the zero point of the chart, around a width of 22 mm and a weight of 57 g. To reach that cloud, the model's line has to carry a big offset, called the bias: a learned constant added so the line does not have to pass through zero.
Centring moves the cloud so it surrounds zero, and that improves two numbers at once:
| | Raw data | Centred data | |---|---|---| | Size of the data (distance of the farthest example from zero) | about 74 | about 13 | | Margin (empty space around the dividing line) | 0.045 | 0.989 | | Worst-case limit on corrections | 2,633,550 | 168 | | Corrections actually made | 29,870 | 1 | | Passes needed | 11,976 | 2 |
The worst-case limit comes from the convergence theorem, the previous concept in this series: size divided by margin, squared. A smaller size and a wider margin together cut it by a factor of roughly 15,000. In the real run, the centred model fixed its line with a single correction and was done.
A common misconception
"If you subtract numbers from the data, you are changing the data." You are changing where it is drawn, not what it says. Part A is still exactly as much wider than part B as it was before. The model just no longer has to stretch to reach a faraway cloud.
Why it matters
Before, training needed about 60 times more passes than a 200-pass budget allows; after, it was over almost immediately. The learner did not get any smarter. The data was placed where the learner could work with it.
The course treats this as the first case of a pattern that returns later, when choosing a network's starting values and how fast it is allowed to learn: the mathematics says a result is possible, and practical choices like this one decide whether you actually get it in a reasonable time.
Learn it step by step in Chapter 1 of our free course AI From Scratch: The Perceptron From Scratch: What a Neuron Computes.
Aussi sur
Plus de vidéos de ce chapitre
#01
#001Perceptron: the yes-or-no model every neural network grew from
#02Machine learning: how a computer draws its own dividing line
#002Decision boundary: the line a simple model draws between yes and no
#003Dot product: the multiply-and-add step at the heart of AI's math
#004Learning from mistakes: the perceptron rule that only moves when wrong
Texte rédigé avec l’aide de l’IA à partir du chapitre 1 de notre cours gratuit.