This is a complete explainer of multi-head attention in transformers. Its central idea is that a single attention head produces just one softmax distribution per token and therefore one weighted average, one relationship, so to attend to several relationships at once the model runs h attention operations in parallel, each in its own learned subspace of size d_k equal to d_model divided by h, and because the dimension is split rather than multiplied the whole thing costs about the same as one full-width head, after which an output projection W_O mixes the heads back together. It motivates this with an example where representing a verb needs subject-verb agreement with a distant noun, the intervening clause, and the previous token all at once, which one softmax cannot separate, and notes that making a single head wider does not help because it is still one average. Each head i computes Attention of X times W_i^Q, X times W_i^K, X times W_i^V, where W_i^Q and W_i^K are d_model by d_k and W_i^V is d_model by d_v, with d_k equal to d_v equal to d_model over h; in the original transformer d_model is 512, h is 8, and d_k is 64. Each head runs full scaled dot-product attention independently in its subspace, so each learns its own notion of relevance. The head outputs, each n by d_v, are concatenated along the feature dimension back to n by d_model since h times d_v equals d_model, then multiplied by the output projection W_O which is h times d_v by d_model; W_O is not bookkeeping, it lets every output dimension be a learned mixture of all heads and writes the result back into the residual stream. A multi-head attention layer has four projection matrices W_Q, W_K, W_V, and W_O, giving four times d_model squared parameters plus biases. The exact tensor shapes with a batch dimension B are: input X is B by n by d_model, the Q K V projections are B by n by d_model, splitting into heads gives B by h by n by d_k, the scores Q K transpose over root d_k are B by h by n by n, the softmax weights are B by h by n by n, the head outputs A times V are B by h by n by d_v, merging heads returns B by n by d_model, and the output projection keeps B by n by d_model, so the layer is shape preserving which is why blocks stack. On what heads learn: previous-token heads attend from each position to the one before it; positional or adjacency heads attend by fixed offset or to delimiters; induction heads implement the pattern A B then later A predicts B by attending from the second A to the token that followed the first A, and Olsson and colleagues in 2022 argue induction heads are the primary mechanism of in-context learning and emerge in a sharp phase change, formed as a two-head circuit of a previous-token head plus an induction head; syntactic and rare-word heads track dependencies or the least frequent token, per Voita and colleagues in 2019. Caveats: many heads are redundant, and Voita et al pruned 38 of 48 encoder heads for only 0.15 BLEU drop with specialized heads pruned last; head interpretability is partial and post hoc; the KV cache grows with the number of heads which multi-query and grouped-query attention address by sharing key and value heads; and pushing h too high makes d_k too small. It closes by placing multi-head attention inside the full transformer block with residual connections, normalization, and a feed-forward network, and points to grouped-query attention and mechanistic interpretability. Two interactive widgets let the reader set d_model, h, n, and B and see every tensor shape and the parameter count with a divisibility check, and select a head, previous-token, positional, induction, content, or one averaged head, to see its attention pattern on a sentence with a repeat.