#04Este vídeo e o respetivo texto estão em inglês.
Condition number: why a stretched valley makes AI training crawl
A long, narrow error valley forces tiny steps and slow learning. See what the condition number measures and why centring data cut 79,513 steps to 10.
Em detalhe
What the condition number is
Training an AI model is a walk downhill. Picture a landscape where every spot is one possible set of the model's settings and the height is how wrong the model is. Gradient descent, the standard way to train, keeps stepping in whichever direction goes down fastest.
The condition number describes the shape of that landscape near the bottom. It compares how sharply the ground bends in its steepest direction with how gently it bends in its flattest one. A perfectly round bowl scores 1. A long, narrow trench can score in the tens of thousands.
The shape matters for one reason: the walker uses a single step size, called the learning rate, for every direction at once.
What the video shows
The video opens with the hook "same data, same code, 8,000× more work". Its point: when the error landscape is a long, narrow trench instead of a round bowl, learning zigzags and seems to go on forever. Strictly, it does reach the bottom in the end, just after a huge number of steps. Then comes the result from our course: centring the data cut the work from 79,513 steps to 10, roughly 8,000 times less.
An everyday example
Imagine tinting a tin of white paint with two colours, using one spoon for both. The black tint is so strong that one heaped spoonful ruins the tin, so you only dare a tiny pinch per round. The pale blue tint is weak, and you need a lot of it. With pinches small enough for the black, adding enough blue takes hundreds of rounds. One direction punishes big moves, the other needs them, and one step size cannot suit both. That is a badly conditioned problem.
How it works
In a trench, the walls are steep and the floor is nearly flat. The steep walls put a hard limit on the step size: go above it and every step overshoots further than the last, until the numbers explode. So you must pick a step small enough for the walls. Along the gentle floor, that same small step barely gets anywhere.
In the course, parts on a conveyor belt are described by two raw measurements, in millimetres and grams. Those numbers sit far from zero, and the valley they produce bends 33,452 times more sharply one way than the other. The fix is centring: subtracting the average from each measurement so the values sit around zero. Afterwards the condition number is 7.44. Each version got the best step size it could handle:
| Inputs | Condition number | Steps to get within 1% of the best answer | |---|---|---| | Raw millimetres and grams | 33,452 | 79,513 | | Centred | 7.44 | 10 |
Same data, same code, same final answer. Only the shape of the valley changed. Momentum, covered in its own short, attacks the same trench another way.
How to spot it
- Every step size big enough to make quick progress makes training blow up, while with the safe ones the error creeps down for thousands of steps.
- Your inputs sit far from zero, like raw weights and lengths that were never centred.
A common misconception
Normalising inputs, which means putting them on a common scale centred on zero, is often taught as good housekeeping. The course makes a stronger point: it is arithmetic. It changes the shape of the valley, and the shape decides how many steps you need.
Why it matters
A badly conditioned problem does not fail loudly. Training still runs; it is just thousands of times slower than it needs to be, and that is easy to blame on the computer or the model. Here, subtracting an average from each input was the difference between 10 steps and almost 80,000.
Learn it step by step in Chapter 3 of our free course AI From Scratch: Downhill: Gradient Descent, and the Two Steps Everyone Skips.
Também em
Mais deste capítulo
#04
#021Grid search: why you cannot train an AI by trying every setting
#022Derivative: how much the result moves when you nudge one input
#023Chain rule: how AI traces the effect of every layer
#024Gradient: the arrow that points uphill, and why AI walks the other way
#025Gradient descent: how AI finds its way downhill in small steps
Texto escrito com assistência de IA a partir do capítulo 3 do nosso curso gratuito.