This is an explainer of Multi-Query Attention from Shazeer 2019 and Grouped-Query Attention from Ainslie and colleagues 2023. Its central idea is that autoregressive decoding is bottlenecked not by attention’s FLOPs but by streaming the KV cache from memory every single token, and the cache grows with the number of key and value heads, so since a model needs many distinct query heads but not many distinct key and value heads, you share them: Multi-Query Attention keeps one shared K and V head, and Grouped-Query Attention keeps one K and V head per group. As the quadratic-wall post showed, decoding is memory-bandwidth-bound because producing each new token means re-reading the entire KV cache of all past tokens’ keys and values, so the cache not the arithmetic is the bottleneck. In standard multi-head attention every one of the h heads stores its own keys and values for every past token, so the KV cache is two times n times d_model times L bytes where d_model equals h times d_head, scaling with the number of KV heads, while queries are recomputed fresh each step and never cached, and that asymmetry is the opening. Multi-Head keeps h query, h key, h value heads, maximum quality and cache. Multi-Query, Shazeer 2019, keeps h query heads but collapses to one shared key head and one shared value head so every query head attends against the same K and V, shrinking the cache by a factor of h, for a 32-head model 32 times less memory to stream each decode step which translates almost directly into faster generation, at the cost of some quality loss and sometimes training instability. Grouped-Query, Ainslie et al 2023, is the dial between them: partition the h query heads into G groups each with its own shared K and V head, so G equals h is MHA, G equals 1 is MQA, and the KV cache reduction is h over G, with a middle G often 8 keeping most of MHA’s quality at close to MQA’s cache. Intuitively the KV heads encode what information each past token makes available while query heads encode what each position wants to look up, and you need diversity mostly on the asking side, so a handful of shared KV projections expose enough while distinct query heads preserve varied attention. A key practical result is uptraining: you can convert an existing multi-head checkpoint to GQA by mean-pooling its K and V heads into the target number of groups then fine-tuning for around 5 percent of the original pretraining compute, recovering most of the lost quality, which is why GQA spread fast since labs could retrofit existing models. Caveats: quality versus cache is a real trade and MQA can measurably degrade quality and destabilize training which is why GQA exists; it does not touch prefill or the order n squared compute, only the decode-time KV cache; and the cache still grows linearly with context so GQA cuts the constant factor not the growth, and pushing further means compressing the KV not just sharing it. In the lineage MQA and GQA are the KV-cache-reduction branch leaving the attention math exact and shrinking the decode state; the line continues with Multi-head Latent Attention from DeepSeek which projects keys and values into a small shared latent that is cached, and the most radical branch, linear attention and state-space models, replaces the growing cache with a fixed-size recurrent state. Adopted in Llama 2 70B with 64 query heads and 8 groups, Mistral, and Falcon. Two interactive widgets let the reader drag the number of groups G to watch 8 query heads regroup onto shared K and V heads with the cache and reduction updating, and compare KV cache growth of MHA, GQA-8, and MQA as context length increases.