Sari la conținut

Acest videoclip și textul său sunt în engleză.

Gradient checking: catching the silent bug in AI training math

AI Concepts #032

A hand-derived slope formula can be quietly wrong with no error message. Learn how gradient checking measures slopes directly and exposes a one-character typo.

În detaliu

What gradient checking is

To train an AI model, you need its gradient: the set of slopes that says, for each setting of the model, how much the error changes if you nudge that setting a little. Training follows those slopes downhill, step after step.

Usually someone works out a formula for those slopes, by hand or in code. And formulas can be wrong. The nasty part is that a wrong gradient does not crash anything. The program runs, no error message appears, and the mistake stays invisible.

Gradient checking is the safety net. Instead of trusting the formula, you measure the slopes directly and compare the two.

What the video shows

The video makes the point in two beats. First: calculated slopes can be silently wrong. Second: measure them with tiny nudges and compare. In the course, a single missing character in the formula jumped out immediately.

An everyday example

Imagine you doubt your phone's step counter. It works out your steps with a formula you never see. You do not need to open the app's code: walk exactly 100 steps, counting them yourself, then look at the screen. If it says 99 or 101, it is fine. If it says 67, something in its formula is broken, even though the app never showed an error. Your slow, simple count is the measurement; the app's number is the formula. Gradient checking is the same test for a model's slopes.

How it works

The measuring trick takes three moves for each setting:

  1. Nudge the setting up by a tiny amount, such as 0.00001, and record the error.
  2. Nudge it down by the same amount and record the error again.
  3. Take the difference between the two errors and divide it by twice the nudge. That is the measured slope.

Nudging both ways, rather than only upwards, is called a central difference. For the same size of nudge it is far more accurate, because the biggest source of measurement error cancels out. Then you compare every measured slope with the one your formula produced.

How to read the result

The comparison uses a relative error: the size of the gap divided by the size of the slopes themselves. A raw gap tells you little on its own. A gap of 0.0001 would be a disaster on a slope of 0.001, and meaningless on a slope of a million.

Here is what the course found:

| Slope formula | Relative error | |---|---| | Correct, worked out by hand | about 2 parts in 100 billion | | Same formula, with one factor of 2 left out | about one third |

The rule of thumb:

  • below roughly one part in ten million means the formula and the measurement agree;
  • above roughly one part in ten thousand means there is a bug.

A one-character typo turned a near-perfect match into a gap of a third. No error message would ever have shown it.

Why it matters

Hand-derived slope formulas are exactly where small typos hide, and a wrong slope does not announce itself: training just follows it. That makes the check more than a nice extra. The course keeps it for later, and in Chapter 5 it is used to debug an engine that computes gradients automatically. Without a check like this, a wrong gradient would be almost impossible to find.

Learn it step by step in Chapter 3 of our free course AI From Scratch: Downhill: Gradient Descent, and the Two Steps Everyone Skips.

Și pe

Mai multe din acest capitol

Text scris cu asistență AI din capitolul 3 al cursului nostru gratuit.