This is a from-scratch explainer of causal masking in decoder-only transformers. Its central idea is that a causal mask forbids every position from attending to any position to its right by setting those attention scores to negative infinity before the softmax, which makes a single parallel forward pass compute at every position exactly the prediction it would make if it had only seen the past, so that training on all positions at once with true tokens and inference one token at a time on the model’s own tokens become the same function. It begins with the leak: because self-attention lets position i attend to every position including i plus one, and the training target at position i is token i plus one, an unmasked autoregressive model learns the trivial shortcut of copying the next token through the attention path, driving training loss to near zero while being useless at generation where the future does not yet exist; this is label leakage. The fix adds a mask matrix M to the scores before the softmax, where M is zero for positions j less than or equal to i (present or past) and negative infinity for j greater than i (future), giving masked attention as softmax of Q K transpose plus M over the square root of d k, times V. Because softmax exponentiates each score and the exponential of negative infinity is zero, every future position contributes exactly zero weight, and the softmax denominator sums only over surviving past and present entries so each row renormalizes to one over the tokens it is allowed to see; the mask is a fixed lower-triangular matrix, in real code a large negative constant rather than literal negative infinity. Masking before rather than after the softmax matters because zeroing weights after softmax leaves them summing to less than one, an un-normalized leaky distribution. The payoff is that with the future masked, one forward pass computes the correct next-token prediction for every position simultaneously and independently, so training is just shifting the sequence by one to make inputs x1 to x T minus 1 and targets x2 to x T, running one masked pass, and averaging cross-entropy; feeding the true previous tokens rather than the model’s own predictions is teacher forcing, from Williams and Zipser 1989, which is what lets all positions be computed at once. Training versus inference is the same mask in two worlds: at training the whole target is known so teacher forcing feeds real tokens and the mask enforces a rule, while at inference the future does not exist so the model generates token by token appending each output and the mask enforces a fact. The seam is exposure bias, named by Bengio, Vinyals, Jaitly, and Shazeer 2015, the train-test mismatch where the model trains only on correct prefixes but at inference conditions on its own possibly wrong prefixes so errors compound; scheduled sampling and sequence-level objectives mitigate it, and this is one reason low teacher-forced perplexity does not guarantee good free-running generation. The mask is also the entire difference between an encoder, which is unmasked and bidirectional like BERT and good at understanding but unable to generate, and a causal decoder like GPT that can generate left to right. Costs: masking does not save the quadratic compute though kernels like FlashAttention skip fully masked blocks, it enables the KV cache because earlier tokens’ keys and values never change, it forfeits right context for generation, and teacher forcing differs from free-running. It closes by pointing to the KV cache and to decoding and sampling. Two interactive widgets let the reader toggle a causal mask on a six-token attention matrix and watch the upper triangle go to zero while rows renormalize, and switch between parallel teacher-forced training and sequential autoregressive inference.