#04Dieses Video und der Text sind auf Englisch.
Grid search: why you cannot train an AI by trying every setting
Trying every combination finds the best settings for a tiny model. Here is why each extra setting multiplies the work until search becomes impossible.
Im Detail
What grid search is
A model is a formula with adjustable numbers inside it, called parameters. Training means finding the values of those numbers that make the model's mistakes as small as possible. Grid search is the most obvious way to do that: write down a list of candidate values for each number, try every combination, measure the error each time, and keep the winner.
It is honest and simple, and for a model with two numbers it works. The trouble starts when you add a third, a tenth or a twenty-thousandth.
What the video shows
With two settings, trying every combination is fine. Each new setting multiplies the number of options, and by the time you reach a small neural network there are more combinations to try than there are atoms in the universe.
An everyday example
Imagine a bike lock with a single dial of ten digits. You can try all ten in a few seconds. Add a second dial and there are 100 combinations. A third dial makes 1,000, a fourth 10,000. Each time you only added one small dial, but the work multiplied by ten.
Grid search behaves exactly like that lock. Every parameter of a model is one more dial, and each dial has far more than ten positions.
How it plays out in the course
The course starts with eight parts from a factory conveyor belt and asks a simple question: can a straight line predict a part's weight from its width? A line has just two numbers: its slope (how steep it is) and its offset (how high it sits).
Try every slope from 0 to 5 and every offset from -5 to 5, in steps of 0.01. That is 501 × 1,001 = 501,501 tries. It works: the best line has a slope of 2.1 and an offset of 0. But it took half a million evaluations to pin down two numbers to two decimal places. Later in the same chapter, gradient descent reaches four decimal places in eight steps.
Now keep 1,000 candidate values per parameter and watch what happens to the count:
| Model | Parameters | Combinations to try | |---|---|---| | The straight line | 2 | 10^6, a million | | A tiny network that solves XOR, a classic "one or the other, but not both" puzzle | 9 | 10^27, a 1 followed by 27 zeros | | A small multi-layer network | 20,000 | 10^60,000, a 1 followed by 60,000 zeros |
For scale, the observable universe holds roughly 10^80 atoms.
Why a faster computer does not save you
It is tempting to treat this as a speed problem. It is not. With 1,000 values per parameter, a computer that is a thousand times faster buys you exactly one extra parameter, and every setting you add multiplies the bill again. So search does not just become slow as models grow. Past a handful of parameters, it is no longer an option at all.
Why it matters
That table is the reason AI models are trained the way they are. Instead of checking every option, training measures which way the error goes down and takes small steps in that direction, which is the idea behind gradient descent. Grid search remains a fair tool when only a couple of numbers need setting, and the course uses it to find its best line. For anything that deserves the name neural network, it is off the table.
Learn it step by step in Chapter 3 of our free course AI From Scratch: Downhill: Gradient Descent, and the Two Steps Everyone Skips.
Auch auf
Mehr aus diesem Kapitel
#04
#022Derivative: how much the result moves when you nudge one input
#023Chain rule: how AI traces the effect of every layer
#024Gradient: the arrow that points uphill, and why AI walks the other way
#025Gradient descent: how AI finds its way downhill in small steps
#026Learning rate: the step size with a ceiling you can calculate
Text mit AI-Unterstützung aus Kapitel 3 unseres kostenlosen Kurses geschrieben.