This is an interactive deep dive on the claim that “compression is intelligence,” building the mathematics of information theory from scratch and connecting it to how large language models are trained. It opens with the Hutter Prize, which offers 500,000 euros for compressing a one-gigabyte snapshot of English Wikipedia (the file enwik9) and frames text compression as a path to artificial general intelligence, because to compress text you must predict it, and to predict it you must model the world. The core question is the minimum number of bits needed to send a message. A fixed-length code such as ASCII (8 bits per character) or a 2-bit rover instruction code wastes bits when symbols are not equally likely. A moon-rover example with instruction probabilities one-half, one-quarter, one-eighth, one-eighth motivates variable-length prefix codes and the Kraft inequality, which says the codeword-length budget sum of two-to-the-minus-length must be at most one. Minimizing expected code length subject to that budget, via a Lagrange multiplier, yields the central formula: the optimal code length for an outcome of probability p is negative log base two of p, called self-information or surprisal. A bit is a halving of possibility; a certain event carries zero bits, a coin flip one bit, a one-in-1024 event ten bits. Because independent probabilities multiply and logarithms turn products into sums, the surprisals of a message’s symbols add. The average surprisal of a distribution is Shannon entropy, H(p) equals minus the sum of p log p, which Shannon’s source coding theorem proves is the hard floor on lossless compression, the average bits per symbol no code can beat. Entropy is maximal for a uniform distribution and approaches zero for a predictable one: predictability is compressibility. Huffman coding rounds code lengths to integers and stays within one bit of entropy, while arithmetic coding encodes an entire message as a single number in the interval zero to one, subdividing the interval by symbol probabilities so the final interval width is the product of probabilities and the code length is the sum of surprisals, achieving fractional bits and coming within about two bits of the entropy of the whole message. Because the true distribution p is unknown, we use a model q; coding p-distributed data with a q-based code costs the cross-entropy H(p,q) equals minus the sum of p log q, which decomposes exactly into the entropy floor H(p) plus the Kullback-Leibler divergence D_KL(p given q), the extra wasted bits from an imperfect model, zero only when q equals p by Gibbs’ inequality. The reveal: a language model’s pre-training objective is cross-entropy loss on next-token prediction, minus the average of log q of each token given its context; each term is exactly the bits arithmetic coding would spend to encode that token using the model’s predicted distribution, so minimizing next-token cross-entropy loss is identical to minimizing the compressed size of the training data. Prediction and compression are two sides of one optimization: any predictor plus an arithmetic coder is a lossless compressor, and any compressor implies a predictor. Shannon estimated the entropy of English at 0.6 to 1.3 bits per character in 1951 using his wife Betty as a human next-letter predictor, far below ASCII’s 8 bits and letter-frequency’s roughly 4 bits. DeepMind’s 2023 paper Language Modeling Is Compression showed Chinchilla 70B, a text model, used as an arithmetic-coding engine compresses ImageNet image patches to 43.4 percent (beating PNG’s 58.5 percent) and LibriSpeech audio to 16.4 percent (beating FLAC’s 30.3 percent), because next-chunk prediction is a universal skill. The honest caveats: those ratios exclude the 70-billion-parameter model needed to decode, and counting the model as part of the description length gives a compression view of scaling laws and Occam’s razor with a dataset-dependent optimal model size; the true compression floor is the Kolmogorov complexity, which is uncomputable; “intelligence” is ill-defined so compression is necessary for prediction but not provably sufficient for intelligence; and tokenization changes the units. The exact, provable core is that a language model’s cross-entropy loss equals the compression rate it achieves, so lower loss means a smaller file and a model that has captured more of the structure of its data.