Skip to lesson content

THE ILLUSTRATED LLM TUTORIAL / 02 OF 15

Tokenization in LLMs

Convert text to token IDs and explain why token counts depend on the tokenizer.

Try the example ↓
Tokenization in LLMs: Text: I love learning!; Pieces: I / love / learn / ing / !; IDs: 1 / 2 / 3 / 4 / 5; Budget: 5 tokens in this toy split
Lesson 02 visual guide · Read the four steps, then explore the explanation below.
  1. 01TextI love learning!
  2. 02PiecesI / love / learn / ing / !
  3. 03IDs1 / 2 / 3 / 4 / 5
  4. 04Budget5 tokens in this toy split

Tokens and vocabulary

A tokenizer maps text into discrete units and then integer IDs in a vocabulary. Units may be words, subwords, characters, or bytes. Spaces and punctuation can affect how text is split.

Why subwords help

A word-only vocabulary struggles with rare or new words. Subword methods can represent a rare word as smaller known pieces. Byte-level coverage can avoid unknown characters, though not every tokenizer uses it.

Algorithms and special tokens

BPE repeatedly merges frequent adjacent pieces during tokenizer training. WordPiece uses a different vocabulary-building criterion. Unigram selects among possible segmentations. SentencePiece is a toolkit supporting algorithms such as BPE and Unigram, not a single competing algorithm. BOS, EOS, PAD, and UNK are model-dependent special tokens.

Budgeting context

Count tokens with the tokenizer used by the actual model. Instructions, chat formatting, retrieved text, conversation history, and output can all consume capacity. Reserve space for the answer instead of filling the input to its limit.

Worked example

We explicitly split learning into learn and ing. These IDs are invented for the lesson; they are not the IDs used by any commercial model.

Download lesson 02 Python example

Python 3 / standard library
vocabulary = {"I": 1, "love": 2, "learn": 3, "ing": 4, "!": 5}
pieces = ["I", "love", "learn", "ing", "!"]
ids = [vocabulary[piece] for piece in pieces]
reverse = {value: key for key, value in vocabulary.items()}
print(ids)
print([reverse[i] for i in ids])
window, instructions, documents, reserve = 100, 10, 55, 20
print("History budget:", window - instructions - documents - reserve)

Expected output

[1, 2, 3, 4, 5]
['I', 'love', 'learn', 'ing', '!']
History budget: 15

Encoding maps pieces to IDs; decoding recovers pieces. Reconstructing exact text also needs the tokenizer's whitespace and byte rules. The budget leaves 15 tokens for history in this simplified example.

Practice and self-check

Common mistake

Do not assume one word equals one token or copy the historical model context limits in the source sheet as universal limits.

Student tasks

  1. Add the token reader and encode a sentence containing it.
  2. Explain what happens if an input piece is missing from this vocabulary.
  3. For a 500-token window with 60 instruction tokens, 250 document tokens, and 100 output tokens, calculate the history budget.
Checkpoint — open after attempting the tasks

The history budget is 90 tokens. A real tokenizer may split an unknown word; this dictionary example raises a KeyError.