This is an explainer of Mamba and state-space models by Gu and Dao 2023. Its central idea is that attention and state-space models are opposite answers to how you remember the past: attention keeps a perfect transcript, the KV cache, order n state with perfect recall, and re-reads all of it every step so it costs order n squared; a state-space model compresses all history into a fixed-size state, order one memory per step and order n compute with no growing cache, but must forget. Classic state-space models lost to attention on language because their fixed dynamics forgot indiscriminately, and Mamba’s insight is to make the state update input-dependent, selective, so the model decides what to remember or forget based on the token it is reading, recovering content-based memory at linear cost. An SSM maps input to output through a hidden state with a linear update where the new state equals A times the previous state plus B times the input and a linear readout where the output equals C times the state; it is a linear RNN whose fixed-size state carries a compressed summary of everything seen so far, so inference is order one per step in memory with no cache that grows with context, the property that makes SSMs attractive for very long sequences. Because the recurrence is linear and in classic SSMs time-invariant with A, B, C constant, unrolling it gives a convolution, so the output is the input convolved with a fixed kernel, which parallelizes across the whole sequence like a transformer, giving structured SSMs like S4 the best of both worlds, recurrent at inference with order one state and convolutional at training in parallel order n. The flaw is that fixed A, B, C apply the same dynamics to every token regardless of content, so the model cannot say this token is important remember it, or this is filler ignore it, or reset a new topic starts, which is fine for smooth signals like audio but fatal for language which is about selectively remembering the right earlier token, the content-addressable memory that made attention win, so time-invariant SSMs underperformed transformers on text. Mamba fixes it with a selection mechanism that makes the step size Delta and the matrices B and C functions of the input x_t, so the state update is content-dependent and the model gates information into and out of the state based on what it is reading, selectively remembering ignoring or resetting. The catch is that input-dependent Delta, B, C make the recurrence no longer time-invariant so the convolution trick with its fixed kernel is gone, and Mamba instead computes the input-varying linear recurrence with a hardware-aware parallel scan, an associative-scan algorithm fused into a GPU kernel that keeps the state in fast SRAM and recomputes intermediates in the backward pass, the same IO-aware discipline as FlashAttention, so it stays parallel and fast despite giving up the convolution. The payoff is linear scaling in sequence length, a constant-size state at inference giving roughly 5 times higher generation throughput than a comparable transformer since there is no growing KV cache to stream, and matching or beating transformers of the same size on language, audio, and genomics while handling sequences up to a million tokens, the first SSM to genuinely rival attention on language. Caveats: a fixed state is still lossy so even selective an order one state cannot store arbitrary detail about every past token the way attention’s cache can, and on tasks demanding exact retrieval of a specific earlier token like associative recall and precise copying attention still has an edge; hence many strong models interleave attention and SSM layers, a few attention layers for sharp recall and many Mamba layers for cheap long-range mixing; and the selective scan needs a bespoke hardware-aware kernel, not a drop-in. In the lineage Mamba is the abandon-attention branch and is closer to linear attention than it looks, since a linear-attention model’s running key-value summary is a fixed-size recurrent state, and Mamba-2 in 2024 made this precise showing a duality between selective SSMs and a form of linear attention, so the two most radical efficiency lines converge on the same object, a fixed-size state you update and read instead of a cache you grow and re-scan, which is where the frontier lives, with Kimi K3’s efficient core being a linear-attention mechanism, Kimi Delta Attention, in this same family typically hybridized with full-attention layers to keep the recall a pure fixed state gives up. Two interactive widgets present the attention-versus-SSM tradeoff, order n versus order one state and order n squared versus order n compute, and a selective-versus-fixed state stepper where one signal token must be remembered through filler, showing a fixed recurrence blur and decay the signal while a selective one gates and retains it.