#04Tämä video ja sen teksti ovat englanniksi.
Chain rule: how AI traces the effect of every layer
When one change drives another, which drives a third, the rates multiply. Learn the chain rule, the idea that lets AI trace how each layer shapes the result.
Tarkemmin
What the chain rule is
Many things happen in chains: one change causes a second, which causes a third. The chain rule is the math for chains like that, and it says something simple. The rates multiply. If the first link turns every unit of change into 3 units, and the second link turns every unit into 2, then the whole chain turns every unit into 2 × 3 = 6.
Mathematicians call such a chain a composition: you feed the output of one function (a rule that turns an input into an output) into the next one.
What the video shows
The video frames it as a rule about gears. If one thing speeds up another, which speeds up a third, the speed-ups multiply. That same rule is what lets an AI work out how every layer of a neural network affects the final result.
An everyday example
Unit conversions are chains too. One kilometre is 1,000 metres, and one metre is 100 centimetres, so one kilometre is 1,000 × 100 = 100,000 centimetres. Nobody had to lay a tape measure along a whole kilometre. You multiplied the rate of each link.
Or imagine a small shop where every extra euro spent on flyers brings 3 extra visitors, and every extra visitor brings 2 extra euros of sales. You can tell straight away that each flyer euro brings 6 euros of sales, without following a single customer around.
How it works inside a neural network
A neural network is not merely similar to a chain. It is one. Each layer takes the previous layer's output as its input, and the depth of a network is simply how many links its chain has.
To improve, a network needs to know how much each internal setting affects the final error, the single number that says how wrong it is. Many of those settings sit deep inside the chain, far away from that error. The chain rule solves this: multiply the rates of every link between a setting and the error, and you have that setting's effect.
The course first does this on a simple line that predicts a part's weight from its width. The line's slope changes each prediction, each prediction changes the squared mistake for that part, and the squared mistakes are averaged into the error. Multiply the rates along that short chain and you get exactly how the error responds to the slope. Working this out for one setting while holding all the others still is called a partial derivative. Put them all side by side and you have the gradient.
Why long chains cause trouble
Because rates multiply, a signal travelling back through ten layers gets multiplied by ten numbers. If each of them is a little below 1, say 0.9, only about 0.35 of the signal survives ten layers, and it keeps fading with depth. If each is a little above 1, it grows instead. Those two effects, known as vanishing and exploding gradients, get a whole section later in the course, and they come straight from this one rule.
Why it matters
Everything else in the course rests on this idea. It turns a question that looks hopeless, "how does this one number, buried in layer three, change the final answer?", into a chain of multiplications that a computer can do in a blink.
Learn it step by step in Chapter 3 of our free course AI From Scratch: Downhill: Gradient Descent, and the Two Steps Everyone Skips.
Myös palvelussa
Lisää tästä luvusta
#04
#021Grid search: why you cannot train an AI by trying every setting
#022Derivative: how much the result moves when you nudge one input
#024Gradient: the arrow that points uphill, and why AI walks the other way
#025Gradient descent: how AI finds its way downhill in small steps
#026Learning rate: the step size with a ceiling you can calculate
Teksti on kirjoitettu AI-avusteisesti ilmaisen kurssimme luvun 3 pohjalta.