#009Este vídeo e seu texto estão em inglês.
Floating-point numbers: why your computer cannot hold every number
A computer number has a fixed budget of bits, split between detail and range. Learn how that split works and why two 16-bit formats fail so differently.
Em detalhes
What a floating-point number is
A computer gives every number a fixed budget of bits (the 0s and 1s it works with) and splits that budget into three parts: a sign that says plus or minus, the significant digits, called the mantissa, and a scale that says where the decimal point sits, called the exponent. The layout is fixed by an industry standard called IEEE 754.
That split is the whole design. Bits given to the mantissa buy precision: how finely two nearby numbers can be told apart. Bits given to the exponent buy range: how huge or how tiny a number can get before the computer runs out of room. With a fixed budget, you cannot have unlimited amounts of both.
What the video shows
The video explains that a computer stores each number in a fixed number of bits, shared between detail and range. It then compares two 16-bit formats that divide those bits in different ways. One of them can turn tiny values into zero. The other cannot.
An everyday example
Imagine a pocket calculator with an eight-digit screen. It can show 12,345,678 exactly, but a number like 123,456,789,012 does not fit. So it switches to a short form such as 1.2345679E11: the first eight digits, plus a note that says "move the point eleven places". It keeps the size of the number and rounds away the last digits. That is floating point in a nutshell: a limited set of digits plus a point that can move.
Three formats, three different splits
The course compares three formats:
| Format | Total bits | Exponent bits (range) | Mantissa bits (detail) | Largest value | |---|---|---|---|---| | float32 | 32 | 8 | 23 | about 3.4 × 10^38 | | float16 | 16 | 5 | 10 | 65,504 | | bfloat16 | 16 | 8 | 7 | about 3.4 × 10^38 |
Here 10^38 means a 1 followed by 38 zeros. Both 16-bit formats use half the memory of float32, but they spend the savings differently:
- float16 keeps more digits and gives up range. A tiny value like 0.00000001 becomes exactly 0, and 262,016 becomes infinity.
- bfloat16 borrows the whole exponent of float32, so both of those values stay in range (262,016 comes out as 262,144: close, not exact). The price is fewer digits: it stores 0.3 as 0.30078125, an error about sixteen times bigger than the one float16 makes.
Why it matters for AI
The split decides what breaks.
A model learns from gradients: correction signals that say which way, and how much, to nudge each of its settings. Many are tiny, and in float16 a tiny one can round to exactly zero, so the model stops learning from it. That is why float16 training needs a workaround called loss scaling: multiply the loss, the model's error score, by a big constant, so the small signals grow back into the range the format can hold. bfloat16 mostly does not need it.
Even 64-bit numbers have limits. In the course, multiplying together how plausible each of 2,000 measurements is gives infinity, and with a noisier measuring tool the same code gives 0.0. Both answers are wrong, and neither stops the program with an error. In the first case the true value is about 10^608, while the largest 64-bit number is about 1.8 × 10^308.
A common misconception
It is tempting to think computers are exact with numbers. They are exact only with the numbers they can represent. Everything else is rounded to the nearest one they have, and which numbers exist in each format shapes how AI models are trained and stored.
Learn it step by step in Chapter 2 of our free course AI From Scratch: Where a Loss Function Comes From: Likelihood, Not Convention.
Também em
Mais deste capítulo
#009
#010Loss landscape: training an AI means finding the bottom of a valley
#011Chain rule of probability: how AI scores a sentence word by word
#012Bayes' rule: reasoning backwards from a clue to its most likely cause
#013Maximum likelihood: the answer that makes your data least surprising
#014Negative log-likelihood: turning a fragile product into a calm sum
Texto escrito com assistência de IA a partir do capítulo 2 do nosso curso gratuito.