Akshath Tiwari

Language models generate text one token at a time, and each token is a full, memory-bound forward pass through billions of parameters. Speculative decoding gets several tokens per pass without changing the output: a small, fast draft model guesses the next few tokens, and the big target model verifies them all in one parallel pass, keeping the good ones. The trick is a memory-bandwidth loophole (scoring many tokens in a pass costs almost the same as one) plus an exact accept/reject rule: keep a guess with probability min(1, p/q), and on rejection resample from the normalized leftover max(0, p minus q), which makes the output identical to sampling from the target model alone. This interactive essay walks through the bottleneck, the loophole, the draft-verify-settle loop, the exactness proof, the speedup ledger (words per pass and the break-even against drafting cost), and the family of variants (self-speculation, feature-level drafting, n-gram lookahead).