#04Este vídeo y su texto están en inglés.
Gradient: the arrow that points uphill, and why AI walks the other way
The gradient combines the slope of every setting into one arrow. Learn why it points where the error rises fastest, and why AI steps the opposite way.
En detalle
What the gradient is
A model has many adjustable numbers inside it, called parameters. For each one you can measure a slope: if I nudge only this number, how fast does the error change? The gradient is simply all those slopes written together as one list. With two parameters it is a pair of numbers, and a pair of numbers can be drawn as an arrow on a map.
The surprising part is that this arrow is more than bookkeeping. It points somewhere very specific.
What the video shows
The video combines the slope of every setting into one arrow. That arrow points to where the error grows fastest. Turn around and face the opposite way, and you are looking at the fastest way to improve.
An everyday example
Picture a hiking map with contour lines, the curved lines that join points of equal height. Walk along a contour line and you neither climb nor descend. The steepest way up always cuts straight across the contour lines, at a right angle to them.
The gradient is that straight-up-the-hill direction, except the hill is made of error instead of rock. Walk along the arrow and the error climbs as fast as it possibly can. Start walking sideways, at a right angle to it, and at first the error does not change at all. Walk exactly the opposite way and the error falls as fast as it possibly can.
How the course proves it
In the course, a straight line predicts a factory part's weight from its width. The line has two settings: its slope and its offset. At one starting point, the gradient is the pair (-16.385, 8.0). Read plainly: raising the slope would lower the error, and raising the offset would raise it. Drawn as an arrow, it has a length of about 18.23 and points at roughly 154 degrees.
Then comes the test. A search that knows nothing about gradients tries 3,600 directions, one every tenth of a degree, and measures how quickly the error climbs in each one. The steepest climb it finds is at 154.0 degrees: the arrow's own direction, as closely as steps of a tenth of a degree can show. The climb rate there is 18.2337, which matches the arrow's length to six figures.
Why the arrow works
How fast the error changes in a given direction depends on how closely that direction lines up with the arrow. Face along the arrow and you get the full climb. Turn away and the climb shrinks. At a right angle it drops to zero. Face the opposite way and the error falls at the full rate. No direction can beat the arrow itself, which is also why its length equals the steepest possible slope.
Why it matters
This is the real reason for the minus sign in gradient descent, the method that uses this arrow to train models: it steps against the gradient because that is, provably, the direction where error falls fastest. Nobody picked it as a convention. And a model with millions of parameters works the same way. Its arrow simply has millions of parts instead of two, one slope for every setting.
Learn it step by step in Chapter 3 of our free course AI From Scratch: Downhill: Gradient Descent, and the Two Steps Everyone Skips.
También en
Más de este capítulo
#04
#021Grid search: why you cannot train an AI by trying every setting
#022Derivative: how much the result moves when you nudge one input
#023Chain rule: how AI traces the effect of every layer
#025Gradient descent: how AI finds its way downhill in small steps
#026Learning rate: the step size with a ceiling you can calculate
Texto escrito con ayuda de IA a partir del capítulo 3 de nuestro curso gratis.