Akshath Tiwari

This is a from-scratch, derivation-first explainer of language modeling as next-token prediction. It builds one instrument, a surprise meter (surprisal equals negative log probability), and shows how the chain rule of probability, the training loss, and perplexity are three readings of that same instrument. It begins with the naive attempt to model a whole sentence’s joint probability directly, which fails because the number of possible length-N sequences is the vocabulary size raised to the power N (about ten to the ninety-fourth for a fifty-thousand-token vocabulary and twenty tokens). The chain rule of probability rescues this by factorizing the joint probability exactly into a product of conditional next-token probabilities, the autoregressive factorization, turning one impossible question into N ordinary vocabulary-sized classification problems each solved by a softmax over logits. Surprisal, negative log p, is introduced as the per-token score, measured in bits (base two) or nats (natural log). Summing surprisals over a sequence gives the negative log-likelihood, which equals the cross-entropy loss for a one-hot target because every term except the true token vanishes, and minimizing it is maximum likelihood estimation. Averaging over tokens gives the per-token cross-entropy loss; its expectation decomposes into the language’s own entropy plus the Kullback-Leibler divergence between the true and model distributions, so minimizing cross-entropy minimizes distance to the true distribution. Exponentiating the average cross-entropy gives perplexity, the effective branching factor: a perplexity of k means the model is as uncertain as choosing uniformly among k equally likely tokens, with a ceiling at the vocabulary size for a uniform model and a floor at the exponentiated entropy of the language. A fully worked four-token example (the cat sat plus end-of-sequence, with probabilities 0.4, 0.2, 0.1, 0.5) computes a joint probability of 0.004, a negative log-likelihood of 5.52 nats, a cross-entropy of 1.38 nats or 1.99 bits per token, and a perplexity of 3.976, cross-checked three ways. Measured results quote GPT-2 large on WikiText-2: perplexity 19.44 with disjoint 1024-token chunks, 16.44 with a sliding window of stride 512, and about 19.93 reported in the GPT-2 paper. The post closes with the limits of perplexity: it is tokenizer-dependent so it cannot compare models with different tokenizers, low perplexity is not the same as helpfulness or alignment, training uses teacher forcing while generation does not (the exposure gap), and the per-token average can hide a heavy tail of catastrophic tokens. An interactive surprise ledger lets the reader drag the per-token probabilities and watch the joint probability, loss, and perplexity recompute live.