This is an explainer of Multi-head Latent Attention, MLA, introduced by DeepSeek-V2 in 2024. Its central idea is that grouped-query attention shrinks the KV cache by giving many query heads a shared set of key and value heads, which is lossy because fewer distinct key and value heads means less expressive attention and trades a little quality for cache savings, whereas MLA shrinks the cache without giving up per-head diversity by compression instead of sharing: project each token’s hidden state down into a small latent vector, cache only that, and reconstruct the full per-head keys and values from it whenever attention runs. MLA replaces the KV cache with a low-rank latent: down-project the token to a small vector c of dimension d_c much smaller than d_model, cache only c, and up-project it back to full keys and values per head at attention time; because the full head-specific keys and values are reconstructed not shared as in grouped-query attention, head diversity is preserved so the cache shrinks past grouped-query attention and quality stays at multi-head attention level, so compression beats sharing. Mechanically, standard attention derives per-head keys and values and caches all of them, two times number-of-heads times head-dimension numbers per token, but MLA inserts a bottleneck where a shared down-projection compresses the token into a joint latent c equals W_DKV times x of dimension d_c which is cached, and per-head up-projections expand it back to the per-head key and value; only the latent sits in the cache not the full keys and values, the up-projections are model weights not cached with one per head so every head gets its own reconstructed key and value, and the up-projection can be absorbed into the query and output projections at inference so you attend directly in the compact latent space and never materialize full keys and values. The obstacle is that the absorption trick needs the key to be a plain linear function of the cached latent, but rotary position embeddings apply a position-dependent rotation to the key that cannot be absorbed into a query projection because it depends on the relative position of query and key, so low-rank KV compression and RoPE are incompatible out of the box; DeepSeek’s fix is decoupled RoPE, splitting the key into a compressed part from the latent with no RoPE and a small separate part that carries the RoPE rotation shared across heads, with the query getting a matching decoupled RoPE component, so the cache holds the latent c plus that small RoPE key of dimension d_R, still tiny, preserving positional information without breaking compression. The reported numbers beat grouped-query attention on cache while matching multi-head attention on quality, two axes that usually trade off: in DeepSeek-V2 MLA’s cache is roughly 4 to 14 percent of multi-head attention’s, smaller for larger models, and DeepSeek-V3 reports on the order of 70 kilobytes per token versus about 192 to 328 kilobytes per token for grouped-query-attention-based models, a 2.7 to 4.7 times reduction over grouped-query attention, while maintaining performance comparable to and in ablations better than standard multi-head attention, which is why MLA is the KV-cache method of choice for the largest efficient models powering DeepSeek-V2 and the 671 billion parameter DeepSeek-V3. The distinction is precise: grouped-query attention reduces the number of key and value heads, a coarse lossy quantization of head diversity, while MLA reduces the dimension via low-rank projection and reconstructs full per-head keys and values, keeping diversity and only compressing shared information content, and compression dominates. Caveats: extra compute to reconstruct by up-projecting the latent though usually a good trade since decode is memory-bound; the decoupled-RoPE machinery is intricate; and it is an architecture not a retrofit so you cannot cheaply uptrain an existing multi-head-attention checkpoint into MLA the way you can into grouped-query attention. In the lineage MLA is the KV-cache-reduction branch evolved from sharing heads to compressing them, rhyming with Linformer as low-rank compression but along the head and feature dimension rather than the sequence dimension, and combined with mixture-of-experts it is a big part of why frontier models can serve enormous parameter counts at long context, while the frontier’s other move of replacing the growing cache with a fixed-size state is the linear-attention line leading to Kimi K3. Two interactive widgets let the reader set the MLA latent dimension and compare multi-head, grouped-query, and MLA cache sizes per token with the reductions, and step a token through the MLA pipeline from hidden state to down-projected latent plus decoupled RoPE key that get cached to up-projected per-head keys and values that do not, seeing exactly what is stored.