Akshath Tiwari

This is the finale of the attention-efficiency series, an explainer of Kimi Delta Attention, the linear core of Moonshot’s Kimi Linear architecture and the 2.8-trillion-parameter Kimi K3. Its central idea is that linear attention and state-space models keep a fixed-size state with order one memory and order n compute, but their state was lossy because pure linear attention just adds each key-value pair into the state and never removes anything, so associations pile up in superposition and recall degrades, which is why linear models lost to attention on tasks needing precise retrieval; the fix is to make the fixed state editable using the delta rule, an old idea from associative memory that updates a stored association by correcting its error rather than piling new on top of old. Treat the state S as a key-value memory: to store value v for key k, first read what is currently stored, S transpose times k, compute the error v minus S transpose k, and update S proportionally to that error, erasing the stale association and writing the correction; that is the delta rule from Widrow and Hoff, adapted to linear transformers as DeltaNet, and it turns a fixed-size linear-attention state into a genuinely editable memory that fixes the recall weakness. Concretely, pure linear attention builds its state by summing outer products S_t equals the previous state plus phi of k_t times v_t transpose, so writing a new value for a previously seen key just adds on top of the old one and the memory returns a blur of both, whereas the delta rule subtracts before it writes: S_t equals the previous state plus beta_t times the quantity v_t minus the previous state transpose k_t, times k_t transpose, where the previous state transpose k_t is what the memory currently returns for the key so subtracting it removes the stale entry, and if the key is new the error is just v_t and it behaves like ordinary linear attention while if the key was seen before it overwrites cleanly instead of superposing. Gated DeltaNet by Yang and colleagues 2024, improving Mamba2 with the delta rule, added a forget gate that multiplicatively decays the state, combining gating for rapid adaptive erasure with the delta rule for targeted correction. Kimi Delta Attention sharpens this with channel-wise gating: rather than one forget gate per head, each feature dimension of the state gets its own learned decay rate so some channels hold information a long time and others fade fast, a finer temporal control that the Kimi Linear report shows improves stability and long-context recall, and under the hood it runs a hardware-efficient chunkwise algorithm built on a specialized Diagonal-Plus-Low-Rank state transition so it stays linear-time and GPU-friendly, with update S_t equals Diag of alpha_t times the previous state plus beta_t times the error times k_t transpose where alpha_t is a per-channel forget vector. Even a self-editing state is still compressed, so K3 does not go all-linear but hybridizes: Kimi Linear interleaves KDA with full attention in a fixed ratio of 3 KDA layers to every 1 full-attention layer, using Multi-head Latent Attention for the full layer, where the linear layers carry cheap long-range mixing at order n compute and a constant-size state and the periodic full-attention layer supplies exact arbitrary-distance recall, with 3 to 1 reported as the sweet spot between cost and expressivity. The striking claim is that the hybrid is better not just cheaper: it outperformed a full-attention MLA baseline across short-context, long-context, and reinforcement-learning scaling regimes, higher on MMLU, GSM8K, and coding, and on the RULER long-context benchmark at 128k plus it scored 84.3 versus 81.3 for the pure MLA baseline while using a fraction of the KV cache and running faster at long context. Kimi K3, Moonshot’s open 2.8-trillion-parameter model, is built on this foundation, KDA linear layers interleaved with full attention on a mixture-of-experts backbone combined with Attention Residuals across depth, assembling nearly every idea in the series into one system, though K3 is very recent 2026 so the mechanism is solid while exact configuration details are still settling. The synthesis reads as a catalogue of the series: a fixed-size state from linear attention and SSMs; the delta rule for editable associative memory fixing recall; channel-wise gating as the selection idea made per-channel; a hybrid with MLA for occasional exact recall with a cache shrunk by latent compression; hardware-aware chunkwise kernels in the FlashAttention discipline; and Attention Residuals as attention across depth as well as tokens. Caveats: a compressed state is still compressed so the delta rule and gating narrow but do not close the recall gap, which is why the hybrid keeps full-attention layers; the kernels are bespoke and the design space is churning with decoupled erase-write and new gating schemes appearing fast; and the frontier is very new so comparisons are early and the pure-attention-plus-FlashAttention camp remains competitive, a live contest not a closed verdict. Two interactive widgets let the reader store three associations then overwrite one key and watch pure additive memory return the corrupted sum while the delta rule returns the clean value, and adjust the KDA to MLA interleave ratio to see the cost and recall balance shift.