Saltar para o conteúdo

Este vídeo e o respetivo texto estão em inglês.

Gradient descent: how AI finds its way downhill in small steps

AI Concepts #025

Instead of trying every setting, an AI steps against the slope of its error and repeats. Learn how gradient descent works, and why too big a step backfires.

Em detalhe

What gradient descent is

Gradient descent is a way to train a model without trying every possible setting. It repeats one simple move over and over:

  1. Measure the error, the single number that says how wrong the model currently is.
  2. Work out which way the error rises fastest. That direction is called the gradient.
  3. Nudge every setting a small step the opposite way. The size of that step is set by a number called the learning rate.
  4. Repeat until the error stops going down.

That is the whole method. No map of the landscape, no list of every option. Just the slope right under your feet, one step at a time.

What the video shows

The video contrasts this with trying every setting. Find which way the error rises, step the opposite way, repeat: it reaches the same answer in 8 steps instead of half a million guesses.

To be precise about that comparison: in the course, trying every combination on a grid took 501,501 guesses to pin down the best straight line to two decimal places. Gradient descent got four decimal places in 8 steps, and the full answer, to every digit the computer keeps, in 36.

An everyday example

Imagine a shower you have never used before. You do not test every possible position of the tap. You feel the water, notice it is too hot, turn a little toward cold, feel again, and repeat. Each correction uses only what you feel right now. That is gradient descent: measure, adjust a little in the direction that helps, measure again.

And as with a shower, one huge turn of the tap can take you from scalding straight to freezing.

Why a small step actually helps

"Downhill" is only guaranteed for a vanishingly small move. A real step has a size, and the ground can curve underneath you. Close to any point, the error behaves almost like a straight ramp, and on that ramp the math makes a promise: a small step lowers the error by roughly the learning rate times the gradient's length squared.

The promise only holds for small steps, because the bend that the ramp ignores grows with the square of the step size. The course measures exactly how the promise holds up:

| Learning rate | Share of the promised drop you actually get | |---|---| | 0.0001 | 99.94% | | 0.01 | 94% | | 0.1 | 38% | | 0.2 | none: the error goes up |

At 0.2 the promised drop was about 66, and instead the error rose by about 16. The step pointed downhill, and the error still went up.

A useful side effect

Near the bottom of the valley the ground flattens, so the gradient shrinks, and the steps shrink with it, even though the learning rate never changed. Gradient descent slows down on its own as it arrives. The course's version is about twenty lines of code, and it recovers the best possible line to eight significant figures.

Why it matters

Trying every option stops being possible as soon as a model has more than a handful of settings. Gradient descent sidesteps that wall: it never needs to see every option, only the local slope. The one condition it quietly depends on is that each step is small enough, and how small that must be is the job of the learning rate.

Learn it step by step in Chapter 3 of our free course AI From Scratch: Downhill: Gradient Descent, and the Two Steps Everyone Skips.

Também em

Mais deste capítulo

Texto escrito com assistência de IA a partir do capítulo 3 do nosso curso gratuito.