#04Video này và phần văn bản của video bằng tiếng Anh.
Momentum: how remembering past steps stops AI training zigzagging
In a narrow valley, training zigzags and wastes most of each step. Learn how momentum cancels the zigzag, turning 79,513 steps into 461 with two lines of code.
Chi tiết
What momentum is
Training an AI model means walking downhill on an error landscape, where the height is how wrong the model is. At each step, plain gradient descent looks only at the slope right under its feet and steps that way. It has no memory: every step starts from scratch.
Momentum gives it a memory. Instead of following only the current slope, the walker keeps a running average of the slopes it has seen recently and steps along that average. Think of the difference between a ping-pong ball and a bowling ball: the heavy ball keeps rolling the way it has been going and shrugs off small bumps.
What the video shows
The video shows learning in a narrow valley, zigzagging from side to side and wasting most of each step. Remembering past steps cancels the zigzag and keeps the useful direction. The result from our course: 79,513 steps became 461.
An everyday example
Imagine pushing a shopping trolley down a narrow aisle while a friend keeps nudging it left, then right, then left. An empty trolley swerves with every nudge. A fully loaded one barely notices: the left and right nudges cancel each other out, while your steady push forward keeps adding speed. Momentum turns the model into the loaded trolley.
How it works
In a narrow valley, the walls are steep and the floor slopes only gently. So the steps of plain descent mostly go across the valley, bouncing from wall to wall, while progress along the floor is a slow creep. Why the valley gets so narrow in the first place is the subject of the condition number, which has its own short.
Averaging past slopes changes the picture:
- The sideways part of the slope flips every step: left, right, left. In a running average, those flips cancel out.
- The downhill part along the floor points the same way every step. In a running average, it keeps adding up.
In code, this is two extra lines. You keep a running direction, and at each step you shrink the old direction a little and add the new slope to it. How much of the old direction survives is a setting usually called beta: a beta of 0.9 keeps 90% of it each time, and 0.99 keeps 99%.
The numbers
The course tests it on its hardest case: conveyor-belt measurements left in raw, uncentred units, which make a very narrow valley. Each run uses the best step size plain descent can handle and counts the steps needed to get within 1% of the best answer:
| Memory (beta) | Steps needed | |---|---| | None, plain descent | 79,513 | | 0.9 | 1,609 | | 0.99 | 461 |
That is a 172-fold speed-up for two lines of code.
Why it matters
Momentum is a cheap fix for a costly pattern: training that zigzags instead of moving forward. It does not need more data or a bigger computer, only a small memory of where the last steps were heading. It is also a foundation for later ideas: further on, the course builds Adam, a more advanced training method, on this same mechanism. If you ever see a momentum or beta setting in a training tool, now you know what it does: it lets the steady direction build up and the back-and-forth cancel out.
Learn it step by step in Chapter 3 of our free course AI From Scratch: Downhill: Gradient Descent, and the Two Steps Everyone Skips.
Cũng có trên
Thêm từ chương này
#04
#021Grid search: why you cannot train an AI by trying every setting
#022Derivative: how much the result moves when you nudge one input
#023Chain rule: how AI traces the effect of every layer
#024Gradient: the arrow that points uphill, and why AI walks the other way
#025Gradient descent: how AI finds its way downhill in small steps
Nội dung được viết với sự hỗ trợ của AI từ chương 3 trong khóa học miễn phí của chúng tôi.