Akshath Tiwari

This is a full-depth explainer of tokenization in large language models. Its central idea is that a language model never sees letters; it reads a sequence of integer IDs drawn from a fixed vocabulary decided before training, so a token is the model’s atom of perception and anything below the token, such as individual letters or digits, is invisible unless it is its own token. It opens with the strawberry letter-counting failure, framed as a perception bug in tokenization rather than a reasoning bug. It then examines the two naive granularities and why each fails: character-level tokenization has a tiny complete vocabulary and no out-of-vocabulary tokens but produces very long sequences that waste transformer capacity, while word-level tokenization gives short meaningful sequences but an enormous incomplete vocabulary with an out-of-vocabulary problem and lost morphology. Subword tokenization is the compromise. The Byte-Pair Encoding algorithm is introduced with its training rule: start from characters and repeatedly merge the most frequent adjacent pair of symbols until the vocabulary reaches a target size. A worked example uses the corpus hug times ten, pug times five, pun times twelve, bun times four, hugs times five, learning the merges u g into ug at twenty occurrences, u n into un at sixteen, and h ug into hug at fifteen, matching the Hugging Face NLP course. An interactive trainer runs the real greedy algorithm on that corpus and then tokenizes any word with the learned merges. Byte-level BPE, introduced by GPT-2, starts from the 256 raw byte values instead of characters so that every possible string is representable with no unknown token, giving GPT-2 its vocabulary size of 256 bytes plus 50,000 merges plus one end-of-text token equals 50,257, and uses a pre-tokenization regex so merges never cross word boundaries, which is also why a leading space is part of the following token. Two rival schemes are contrasted: WordPiece, used by BERT, merges the pair that most increases corpus likelihood (pair frequency divided by the product of the parts’ frequencies); and Unigram, used by T5 and ALBERT via SentencePiece, starts from an oversized vocabulary and prunes the least useful tokens by likelihood, is probabilistic, and needs no language-specific pre-tokenization. The vocabulary-size trade-off is quantified: larger vocabularies shorten sequences (helping quadratic attention and context usage) but inflate the embedding and softmax parameters, which equal two times vocabulary size times model dimension, and starve rare tokens of training signal; modern models settle around 100K to 256K. The post then explains five tokenization-caused failures: spelling and letter counting due to letters fused inside tokens; arithmetic due to arbitrary multi-digit token grouping without place value, which is why some tokenizers force single-digit tokens; glitch tokens such as SolidGoldMagikarp, TheNitromeFan, RandomRedditorWithNo, and BuyableInstoreAndOnline, caused by a mismatch between the tokenizer training corpus and the model training data producing under-trained token embeddings, documented by Rumbelow and Watkins in 2023; multilingual inequality where under-represented languages cost several times more tokens and thus higher cost, slower responses, and less context; and prompt-boundary brittleness from trailing spaces changing token sequences. It closes by connecting to token-free byte-level models and to why perplexity, being per-token, cannot be compared across tokenizers.