#009Video dan teks ini dalam bahasa Inggris.
Chain rule of probability: how AI scores a sentence word by word
The chance of a whole sentence is the chance of each word given the words before it, multiplied together. Why this exact rule sits under language models.
Detail
What the chain rule of probability is
The chain rule of probability says you can always break the chance of a whole sequence into a chain of smaller chances: the chance of the first item, times the chance of the second given the first, times the chance of the third given the first two, and so on to the end. "Given" simply means "once you already know".
It works for anything that comes in order: coin flips, measurements, cards drawn from a deck, words in a sentence. And it is not a rough estimate or a design choice. It is an exact identity, true every time, straight from what probability means.
What the video shows
The hook calls it the one math rule every chatbot is built on. The point: the chance of a whole sentence equals the chance of each word, given the words before it, all multiplied together. That is always exact, and language models (the AI systems behind chatbots, which read and write text) are built on it.
An everyday example
Take a shuffled deck of 52 cards and draw two. What is the chance that both are aces?
- The first card is an ace with a chance of 4 in 52.
- Once it is, 51 cards remain and 3 of them are aces, so the second card is an ace with a chance of 3 in 51.
- Multiply the two: 4/52 × 3/51 = 1 in 221.
You never listed every possible pair of cards. You walked through the sequence one step at a time, using what you already knew. That is the chain rule.
Now do the same with a sentence. Imagine a model reading "the cat sat". It asks how likely "the" is as a first word, then how likely "cat" is right after "the", then how likely "sat" is right after "the cat". Multiply those three chances and you have the chance of the whole sentence.
Where it comes from
The chain is built from one simpler piece, the product rule: the chance that two things both happen equals the chance of the first, multiplied by the chance of the second once the first is known. Apply the product rule to two items, then again to add a third, then again for a fourth, and the chain appears. (Write that same product rule in both directions and you get Bayes' rule, which has its own video in this series.)
How language models use it
A language model reads text in small pieces called tokens, each a word or part of a word. At every position it gives a probability to each possible next token, based on everything that came before. The chain rule says that multiplying those next-token chances gives the probability of the entire document. The course is blunt: multiplying those chances is not a clever design choice. It is simply this rule at work, and it could not have been anything else.
A common misconception
It is tempting to think that multiplying word-by-word chances is a shortcut that loses something about the whole sentence. It is not. As long as each step can see everything that came before, the chain loses nothing. What can be wrong are the numbers: the rule is exact, but each next-word chance is the model's own estimate, and the total is only as good as those estimates.
Why it matters
When you hear that a chatbot "predicts the next word", this rule is why that is enough. Scoring text one word at a time, each in the light of what came before, is the same thing as scoring the whole text. The course's Chapter 8, on language models, rests on it.
Learn it step by step in Chapter 2 of our free course AI From Scratch: Where a Loss Function Comes From: Likelihood, Not Convention.
Juga di
Lainnya dari bab ini
#009
#010Loss landscape: training an AI means finding the bottom of a valley
#012Bayes' rule: reasoning backwards from a clue to its most likely cause
#013Maximum likelihood: the answer that makes your data least surprising
#014Negative log-likelihood: turning a fragile product into a calm sum
#015Floating-point numbers: why your computer cannot hold every number
Teks ditulis dengan bantuan AI dari bab 2 kursus gratis kami.