Tokens and vocabulary
A tokenizer maps text into discrete units and then integer IDs in a vocabulary. Units may be words, subwords, characters, or bytes. Spaces and punctuation can affect how text is split.
Why subwords help
A word-only vocabulary struggles with rare or new words. Subword methods can represent a rare word as smaller known pieces. Byte-level coverage can avoid unknown characters, though not every tokenizer uses it.
Algorithms and special tokens
BPE repeatedly merges frequent adjacent pieces during tokenizer training. WordPiece uses a different vocabulary-building criterion. Unigram selects among possible segmentations. SentencePiece is a toolkit supporting algorithms such as BPE and Unigram, not a single competing algorithm. BOS, EOS, PAD, and UNK are model-dependent special tokens.
Budgeting context
Count tokens with the tokenizer used by the actual model. Instructions, chat formatting, retrieved text, conversation history, and output can all consume capacity. Reserve space for the answer instead of filling the input to its limit.
Worked example
We explicitly split learning into learn and ing. These IDs are invented for the lesson; they are not the IDs used by any commercial model.
Download lesson 02 Python example
vocabulary = {"I": 1, "love": 2, "learn": 3, "ing": 4, "!": 5}
pieces = ["I", "love", "learn", "ing", "!"]
ids = [vocabulary[piece] for piece in pieces]
reverse = {value: key for key, value in vocabulary.items()}
print(ids)
print([reverse[i] for i in ids])
window, instructions, documents, reserve = 100, 10, 55, 20
print("History budget:", window - instructions - documents - reserve)Expected output
[1, 2, 3, 4, 5]
['I', 'love', 'learn', 'ing', '!']
History budget: 15Encoding maps pieces to IDs; decoding recovers pieces. Reconstructing exact text also needs the tokenizer's whitespace and byte rules. The budget leaves 15 tokens for history in this simplified example.
Practice and self-check
Common mistake
Do not assume one word equals one token or copy the historical model context limits in the source sheet as universal limits.
Student tasks
- Add the token reader and encode a sentence containing it.
- Explain what happens if an input piece is missing from this vocabulary.
- For a 500-token window with 60 instruction tokens, 250 document tokens, and 100 output tokens, calculate the history budget.
Checkpoint — open after attempting the tasks
The history budget is 90 tokens. A real tokenizer may split an unknown word; this dictionary example raises a KeyError.
