ข้ามไปยังเนื้อหา

วิดีโอนี้และข้อความประกอบเป็นภาษาอังกฤษ

Momentum: how remembering past steps stops AI training zigzagging

AI Concepts #031

In a narrow valley, training zigzags and wastes most of each step. Learn how momentum cancels the zigzag, turning 79,513 steps into 461 with two lines of code.

รายละเอียด

What momentum is

Training an AI model means walking downhill on an error landscape, where the height is how wrong the model is. At each step, plain gradient descent looks only at the slope right under its feet and steps that way. It has no memory: every step starts from scratch.

Momentum gives it a memory. Instead of following only the current slope, the walker keeps a running average of the slopes it has seen recently and steps along that average. Think of the difference between a ping-pong ball and a bowling ball: the heavy ball keeps rolling the way it has been going and shrugs off small bumps.

What the video shows

The video shows learning in a narrow valley, zigzagging from side to side and wasting most of each step. Remembering past steps cancels the zigzag and keeps the useful direction. The result from our course: 79,513 steps became 461.

An everyday example

Imagine pushing a shopping trolley down a narrow aisle while a friend keeps nudging it left, then right, then left. An empty trolley swerves with every nudge. A fully loaded one barely notices: the left and right nudges cancel each other out, while your steady push forward keeps adding speed. Momentum turns the model into the loaded trolley.

How it works

In a narrow valley, the walls are steep and the floor slopes only gently. So the steps of plain descent mostly go across the valley, bouncing from wall to wall, while progress along the floor is a slow creep. Why the valley gets so narrow in the first place is the subject of the condition number, which has its own short.

Averaging past slopes changes the picture:

  • The sideways part of the slope flips every step: left, right, left. In a running average, those flips cancel out.
  • The downhill part along the floor points the same way every step. In a running average, it keeps adding up.

In code, this is two extra lines. You keep a running direction, and at each step you shrink the old direction a little and add the new slope to it. How much of the old direction survives is a setting usually called beta: a beta of 0.9 keeps 90% of it each time, and 0.99 keeps 99%.

The numbers

The course tests it on its hardest case: conveyor-belt measurements left in raw, uncentred units, which make a very narrow valley. Each run uses the best step size plain descent can handle and counts the steps needed to get within 1% of the best answer:

| Memory (beta) | Steps needed | |---|---| | None, plain descent | 79,513 | | 0.9 | 1,609 | | 0.99 | 461 |

That is a 172-fold speed-up for two lines of code.

Why it matters

Momentum is a cheap fix for a costly pattern: training that zigzags instead of moving forward. It does not need more data or a bigger computer, only a small memory of where the last steps were heading. It is also a foundation for later ideas: further on, the course builds Adam, a more advanced training method, on this same mechanism. If you ever see a momentum or beta setting in a training tool, now you know what it does: it lets the steady direction build up and the back-and-forth cancel out.

Learn it step by step in Chapter 3 of our free course AI From Scratch: Downhill: Gradient Descent, and the Two Steps Everyone Skips.

มีบน

เพิ่มเติมจากบทนี้

ข้อความเขียนโดยมี AI ช่วย จากบท 3 ของคอร์สฟรีของเรา