This is an explainer of sliding window attention as used in Mistral 7B by Jiang and colleagues 2023. Its central idea is that letting each token attend only to the last few thousand tokens, a fixed window, looks like a hard limit on context but is not, because stacking layers turns a width W window into an L times W receptive field since information hops one window per layer exactly like a convolutional network, and because no token ever attends past W the KV cache is a fixed-size rolling buffer that never grows no matter how long you generate. Sliding Window Attention means each token attends to itself and the previous W minus one tokens, a window of size W which is 4096 in Mistral 7B, and nothing older, so compute is order n times W, linear in sequence length; this is Longformer’s local half but the insight is what depth does to it. The mechanism: in layer one, token i reads tokens from i minus W plus one to i directly, but the representation of the earliest token in that window has itself already absorbed its own window from the previous layer, so by layer two token i is indirectly influenced by tokens up to two W back, and by layer k up to k times W, so the receptive field after L layers is L times W. For Mistral 7B with 32 layers and window 4096 that is a theoretical span of roughly 131 thousand tokens from a window that only ever looks back 4096 directly, information relaying forward through depth like a bucket brigade. The inference payoff may matter more than the compute: in full attention generating token n means storing and re-reading the keys and values of all previous tokens, a KV cache that grows without bound, the memory-bound bottleneck, but with sliding window attention a token only ever attends to the last W keys so you never need to store more than W of them, and the cache becomes a fixed-size rolling buffer where the key and value for timestep i is written to slot i modulo W, overwriting the entry from i minus W which is now out of every future token’s window anyway, so the cache size is W, constant for any n; you can generate a 200 thousand token document with a 4096 slot cache and the decode memory is flat forever. Caveats: beyond L times W is truly gone since the receptive field is a hard ceiling and unlike full attention there is no path to arbitrarily distant context; reaching far is indirect and lossy because long-range influence arrives through many window hops each of which mixes and compresses, so precise recall of a specific token thousands of positions back is weak compared to full attention’s direct link, which is why later Mistral models mix in full-attention layers; and it is a local approximation, good enough for the mostly-local dependencies of language but not for tasks needing exact sharp long-range retrieval where you want full attention made cheap by FlashAttention or a hybrid. In the lineage sliding window attention is the local window of Longformer and BigBird promoted to the whole mechanism for a decoder language model and paired with its killer inference feature the rolling cache, and in practice it rides alongside grouped-query attention since Mistral 7B uses both, with sliding window capping the cache length at W and grouped-query shrinking its per-token size, and modern models increasingly hybridize with some sliding-window layers for cheap local mixing and some full-attention or linear-attention layers for global recall, a cheap-local-plus-a-little-expensive-global shape that runs right up to Kimi K3. Two interactive widgets let the reader increase the number of layers and watch the reachable span grow as L times W on a toy window of 4, and step tokens through a fixed six-slot rolling buffer written at slot i modulo six that never exceeds six while a full cache grows with every token.