Siirry sisältöön

Tämä video ja sen teksti ovat englanniksi.

Tokens: how AI chops your words into numbered pieces

AI 101 #05

Chatbots never see your words, only numbered pieces called tokens. Learn how text gets split, why long words break apart, and why it changes what you pay.

Tarkemmin

What a token is

A token is the small chunk of text that a language AI actually works with. Before a chatbot such as ChatGPT can do anything with your message, a program called a tokenizer cuts it into pieces: whole words, bits of words and punctuation marks. Each piece is then swapped for a number from a fixed list. From that point on, the model only handles numbers.

What the video shows

The video's captions take it in four steps. AI does not read words the way you do. It chops text into tokens. Long words split into several pieces. Then every token becomes a number, and that row of numbers is what the model receives. The post under the video adds a popular rule of thumb about token size; more on that below.

An everyday example

The course opens its chapter on this topic with one word: "strawberry". In its example, the word becomes three pieces, "str", "aw" and "berry", which turn into three numbers: 496, 675 and 15717. The model never gets ten letters. It gets three numbers.

That explains a famous stumble, the one the chapter is named after: asked how many r's are in "strawberry", a model can get it wrong, because it only ever saw three chunks.

Imagine a box of building bricks: a common word like "the" is one ready-made brick, while a rare word is assembled from several smaller ones.

Why not whole words, or single letters?

There are two simpler options, and both break down:

  • Whole words. The list of possible words would be gigantic, and any word the model never met during training, say a brand-new slang term, would have no number at all. "Word" is not even a clear idea everywhere: Chinese and Japanese are written without gaps between words, and German happily joins nouns into ever-longer compounds.
  • Single letters. Nothing is ever unknown, but texts become four to five times longer than they need to be (a 1,000-word document is roughly 5,000 characters), and longer inputs cost far more to process. A single letter also means almost nothing alone, so the model has to spend effort rebuilding the words.

Tokens sit in the middle. Frequent words get one token each; rare ones are split into pieces. Because the smallest possible pieces are raw bytes (the basic units a computer stores text in), there is no text a model cannot represent.

How the split is decided

Nobody writes the list of pieces by hand. The tokenizer is trained on large amounts of text, much like the model itself. The recipe the course builds is called BPE (byte-pair encoding): start from single bytes, glue the pair of neighbours seen together most often into one new piece with its own number, and repeat. On English text, its very first merge joined an "e" with the space after it.

Why it matters

  • Limits and prices are counted in tokens, not words. In the course, one English paragraph of 164 characters came to 31 tokens, about five characters per token. You will often hear "a token is about three quarters of a word". Treat that as a loose rule of thumb only: the ratio is not stable, so a price per token is not a price per word.
  • Not every language is treated equally. A tokenizer trained mostly on English splits other writing systems into more pieces, so the same message can take more tokens, and cost more, in another language.
  • Some odd mistakes make sense. Models work on the pieces, not on what you see. Even "café" can be stored in two ways that look identical on screen yet split into different tokens.

Learn it step by step in Chapter 7 of our free course AI From Scratch: Build a BPE Tokenizer: Why Your Model Can't Count the R's.

Ruudulla

AI doesn’t read words. It chops text into tokens. Long words split into pieces. Each token becomes a number.

Myös palvelussa

Teksti on kirjoitettu AI-avusteisesti ilmaisen kurssimme luvun 7 pohjalta.