#04Bu video ve metni İngilizcedir.
Derivative: how much the result moves when you nudge one input
A derivative measures how fast a result changes when you nudge an input. Learn how AI uses it, and why a nudge that is too tiny gives a worse answer.
Ayrıntılı olarak
What a derivative is
A derivative answers one practical question: if I change this input by a tiny amount, how much does the result change, per unit of change? That is all it is. It is a rate, like kilometres per hour or euros per kilo.
In AI, the result is usually the error, a single number that says how wrong a model is (often called the loss). The inputs are the model's parameters, the adjustable numbers inside it. For one parameter, the derivative tells you whether nudging it up makes the error bigger or smaller, and how strongly.
What the video shows
The video explains the derivative as a nudge: move an input a little and see how far the result moves. Then it shows a surprise. On a real computer, a nudge that is too tiny gives a worse answer, not a better one.
An everyday example
Imagine you want your speed while driving, but all you have is the distance counter and a clock. You note the distance, wait ten seconds, note it again, and divide the extra distance by the time. That gives your average speed over ten seconds. Measure over one second instead and you get closer to your speed at that exact moment. The speed at a single instant is what those averages settle on as the time gap keeps shrinking.
A derivative is the same recipe. Nudge the input, measure the change, divide by the size of the nudge, then let the nudge get smaller and smaller. The number you settle on is the slope of the curve at that point.
How it works, with real numbers
The course takes a straight line that predicts a factory part's weight from its width, freezes one of its two numbers, and asks how the error reacts when the other one, the slope, is nudged away from 1. The exact answer, worked out with calculus, is -16.385.
- With a nudge of 1, the estimate is -8.94. Far off.
- Make the nudge a hundred times smaller and the estimate's mistake also becomes a hundred times smaller. This repeats with perfect regularity down to a nudge of 0.000001.
- The best result comes at a nudge of 0.00000001, where the estimate is right to six decimal places.
The minus sign carries meaning: raising the slope a little makes the error go down, by about 16 units of error for every unit of change.
Why a smaller nudge can be worse
Keep shrinking and the pattern breaks. Below that best nudge, the estimate gets worse again, and with a nudge of 0.00000000000001 it reads -17.05, wrong in the second digit.
The math is fine. The computer is the problem. It stores every number with a limited number of digits. With a nudge that small, the error before and after the nudge agree in their first ten digits. Subtracting one from the other wipes those digits out and leaves mostly rounding noise, and dividing by a tiny nudge blows that noise up. So there is a sweet spot, and going below it is not more careful. It is less.
A small honest footnote
ReLU is a building block used inside many neural networks: it outputs zero for negative inputs and passes positive inputs through unchanged. Exactly at zero it has no derivative, because the slope is 0 when you approach from the left and 1 from the right. Frameworks such as PyTorch simply pick 0. An input of exactly zero almost never happens, and training works either way.
Why it matters
Training a neural network rests on this one measurement: how much does the error change when I nudge this number? Take one derivative per parameter, put them together, and you get the gradient, the direction that training follows downhill.
Learn it step by step in Chapter 3 of our free course AI From Scratch: Downhill: Gradient Descent, and the Two Steps Everyone Skips.
Şurada da
Bu bölümden daha fazlası
#04
#021Grid search: why you cannot train an AI by trying every setting
#023Chain rule: how AI traces the effect of every layer
#024Gradient: the arrow that points uphill, and why AI walks the other way
#025Gradient descent: how AI finds its way downhill in small steps
#026Learning rate: the step size with a ceiling you can calculate
Metin, ücretsiz kursumuzun 3. bölümünden AI desteğiyle yazıldı.