#009Acest videoclip și textul său sunt în engleză.
Loss landscape: training an AI means finding the bottom of a valley
Score a model at every possible setting and its errors form a landscape. Here is why a smooth valley tells training which way is downhill, and a staircase does not.
În detaliu
What a loss landscape is
Every model has settings it can adjust. For each possible setting you can measure how wrong the model is, using its loss (the single score that sums up all its mistakes). Draw that score for every setting and you get a shape: high ground where the model is badly wrong, low ground where it fits well. That shape is the loss landscape, and training is the search for its lowest point.
What the video shows
The hook is that every AI is secretly searching for the bottom of a valley. The short makes one point: chart how wrong a model is at each of its possible settings and you get a landscape. Training is the hunt for its lowest spot, and a smooth valley tells you which way to head.
An everyday example
Imagine standing on a hillside in thick fog, trying to reach the lowest point of the valley. You cannot see the bottom, but you can feel the ground under your feet. If it tilts, you know which way is down. Take a step that way, feel again, repeat.
Now imagine the ground is a staircase instead: flat steps with sudden drops. On a flat step your feet tell you nothing, however close or far the bottom is.
How it works in the course
The chapter uses a machine-shop example: twenty readings of how wide a part is, taken over the hours since a cutting blade was changed. As the blade wears, the parts slowly get wider. A straight line should describe that growth, and the number that matters is its slope: how many millimetres wider the parts get each hour.
A small trick shrinks the problem first. Shift all the readings so that their averages sit at zero. The best line under squared error then passes through that centre point, so only the slope is left to choose.
The course then tries 601 slopes, from 0 to 0.6, and scores each one by its average squared miss. The winner is 0.293 millimetres per hour. The true rate was 0.300, so twenty noisy readings landed within about 2.3% of it. The more interesting result is the shape of the scores:
| Slope tried | Error score | |---|---| | 0.00 | 0.724 | | 0.12 | 0.259 | | 0.24 | 0.034 | | 0.28 | 0.012 | | 0.32 | 0.016 | | 0.40 | 0.105 | | 0.60 | 0.793 |
Read it from top to bottom: the error falls, bottoms out near 0.29, then climbs again. Turned on its side, that is a valley: a single lowest point, smooth walls on both sides, and a clear downhill direction wherever you stand.
Smooth valleys and staircases
The perceptron (one of the simplest learning machines, from Chapter 1) had no such luck. Its score counted mistakes, and a count moves in whole steps, so its landscape was a staircase with flat stretches that gave no direction at all.
The scoring rule shapes the ground too: squaring the misses is what produced this smooth valley. Scoring by the plain average miss leaves a sharp corner at the bottom, and scoring by the single worst miss leaves flat patches where nudging the line changes nothing.
Why it matters
Trying 601 values works when there is only one setting to choose. Real models have far more settings, and testing every combination quickly becomes impossible. Walking downhill works instead, because a smooth landscape tells you, from wherever you stand, which way is down.
That is where the course goes next: Chapter 3 shows how to reach the bottom without testing every point, and what changes when a landscape has several low spots.
Learn it step by step in Chapter 2 of our free course AI From Scratch: Where a Loss Function Comes From: Likelihood, Not Convention.
Și pe
Mai multe din acest capitol
#009
#011Chain rule of probability: how AI scores a sentence word by word
#012Bayes' rule: reasoning backwards from a clue to its most likely cause
#013Maximum likelihood: the answer that makes your data least surprising
#014Negative log-likelihood: turning a fragile product into a calm sum
#015Floating-point numbers: why your computer cannot hold every number
Text scris cu asistență AI din capitolul 2 al cursului nostru gratuit.