본문으로 건너뛰기

이 영상과 텍스트는 영어로 제공됩니다.

Mini-batches: why AI learns faster from a random handful of data

AI Concepts #030

Why training on a random handful of 32 examples beats using the whole dataset at every step: noisier steps, but 219 times less work to reach the same place.

자세히

What a mini-batch is

Training an AI model is a long series of small corrections. Before each one, the model works out which way to nudge its settings to make fewer mistakes. That direction is called the gradient, and it is really an average: the average of what every single training example says.

The obvious approach computes that average over the entire dataset before every step. With a million examples, that means a million calculations just to move once. A mini-batch is the shortcut: pick a small random handful of examples, such as 32, work out the direction from them alone, take the step, then pick a fresh handful. Training this way is called stochastic gradient descent, or SGD, where stochastic simply means random.

What the video shows

The video shows the switch. Instead of checking every example before each step, the model learns from a random handful of 32. The steps get noisier, but they are far cheaper, and in the course it took 219 times less work to reach the same place.

An everyday example

Imagine checking the salt in a big pot of soup. You do not drink the whole pot. You stir it and taste one spoonful. The spoonful is not a perfect measure, but because the pot is stirred, it is a fair one: it does not lean salty or bland in any systematic way. Taste, adjust, stir, taste again. Each spoonful is a mini-batch.

How it works

A random sample gives a noisy estimate of the full average, but an unbiased one: its errors do not lean in any particular direction, so over many steps they tend to wash out. Lots of cheap, slightly wobbly steps beat a few expensive, exact ones.

The course tests this on 100,000 synthetic parts, meaning data generated by a program rather than measured. It counts how many single-example calculations each method needs to get within 0.1% of the best answer:

| Method | Steps | Example calculations | |---|---|---| | Whole dataset at every step | 7 | 700,000 | | Batches of 32 examples | 100 | 3,200 | | A single example per step | 17,580 | 17,580 |

The whole dataset needs the fewest steps but by far the most work. Divide 700,000 by 3,200 and you get the 219 from the video.

Why not just one example?

Going all the way down to one example per step is the most extreme version, and the original form of the idea, which the course credits to Robbins and Monro. It is not the winner: it needs about five times more work than batches of 32. Two facts pull in opposite directions:

  • Small batches are almost free. The chips used for training process big tables of numbers in one go, so 32 examples cost barely more than one, and they point in a much steadier direction.
  • Big batches give diminishing returns. The wobble only shrinks with the square root of the batch size. To halve it, you need four times as many examples, and four times the arithmetic.

So the best choice sits somewhere between one example and the whole dataset. That trade-off is why every training script has a setting called batch size.

Why it matters

Mini-batches are what make large datasets practical. A model with a million examples does not have to finish a million calculations before it can improve at all: it can start correcting itself after a handful. When you see a batch size in any training setup, this is the dial it controls, trading a little noise for a lot of speed.

Learn it step by step in Chapter 3 of our free course AI From Scratch: Downhill: Gradient Descent, and the Two Steps Everyone Skips.

다른 플랫폼

이 챕터의 다른 영상

무료 강좌 챕터 3의 내용을 바탕으로 AI 도움을 받아 작성한 글입니다.