#009Bu video ve metni İngilizcedir.
Negative log-likelihood: turning a fragile product into a calm sum
Multiplying thousands of chances breaks a computer. Taking logarithms turns the product into a sum with the same winner, and squared error falls out of it.
Ayrıntılı olarak
What negative log-likelihood is
Negative log-likelihood is the likelihood (the score for how plausible your data looks under a candidate answer) rewritten in a form a computer can handle. You take the logarithm of the likelihood and flip its sign. The best answer stays exactly the same, but the arithmetic stops breaking, and "higher is better" becomes "lower is better", like any other error score.
What the video shows
The hook: 2,000 multiplications and the computer gives up. Multiplying thousands of chances breaks a computer: with no warning at all, the answer comes out as infinity or as zero. Taking logs turns the product into a calm sum with the same winner.
An everyday example
You already use the idea behind logarithms when you count zeros. Multiplying a thousand by a million is awkward in your head. But a thousand has three zeros and a million has six, so the answer has nine: a billion. Counting zeros turns multiplication into addition. A logarithm does the same thing for any positive number, not just round ones.
Why a computer needs this: it stores every number in a fixed amount of room. Multiply thousands of plausibility scores and the result can grow too large to store, turning into "infinity", or shrink until it becomes exactly zero. (How computers store numbers has its own videos.) Adding thousands of ordinary numbers causes no trouble.
How it works
- Take the logarithm. The product of all the plausibility scores becomes a sum of their logs. A logarithm always goes up when its input goes up, so the answer with the biggest product also has the biggest sum. The winner does not move.
- Flip the sign. By convention the sum is multiplied by minus one, so that smaller means better, just like an error.
- Clean up. Some pieces of the formula do not depend on the answer you are choosing. Removing them, or scaling everything by a fixed positive number, cannot move the lowest point.
The surprise: squared error falls out
Now plug in the bell-curve assumption from maximum likelihood: each measurement is the true value plus a small random error, where big errors are rare. The bell-curve formula keeps each squared gap tucked inside an exponential, a number raised to a power. Taking the log undoes that power exactly and leaves the squared gap out in the open. Add them up, clean up, and what remains is the sum of squared gaps between prediction and measurement: squared error is a negative log-likelihood in disguise. What that means for anyone who uses it is the subject of the mean squared error video.
The chapter checks this on the same 601 candidate slopes: both curves bottom out at exactly the same slope, 0.293. Their heights differ, and the negative log-likelihood even dips below zero, to about -17, which a sum of squares never can. That is allowed: bell-curve scores measure how tightly values crowd around a point, so they can be bigger than 1, and the negative log of a number above 1 is below zero.
A caution from the course
Throwing pieces away is only safe while they truly do not depend on what you are fitting. If the model also chooses how wide its bell curve is (how much noise it expects), one discarded piece suddenly matters: it is what stops the model claiming zero noise, which would make the data look infinitely plausible.
Why it matters
This is the pattern behind every loss in the course: state the noise you expect, take the negative log, read off the loss. It makes a fragile calculation stable and ties the score you minimise to an assumption you can inspect.
Learn it step by step in Chapter 2 of our free course AI From Scratch: Where a Loss Function Comes From: Likelihood, Not Convention.
Şurada da
Bu bölümden daha fazlası
#009
#010Loss landscape: training an AI means finding the bottom of a valley
#011Chain rule of probability: how AI scores a sentence word by word
#012Bayes' rule: reasoning backwards from a clue to its most likely cause
#013Maximum likelihood: the answer that makes your data least surprising
#015Floating-point numbers: why your computer cannot hold every number
Metin, ücretsiz kursumuzun 2. bölümünden AI desteğiyle yazıldı.