This is an explainer of the 2020 Longformer by Beltagy, Peters, and Cohan and BigBird by Zaheer and colleagues. Its central idea is that a pure sliding window is cheap but cuts the long-range links attention exists for, so Longformer and BigBird add a handful of global tokens that attend to and are attended by everything, restoring full-graph connectivity in two hops at order n cost, and BigBird adds random edges and proves the result is as expressive as full attention. The recipe combines three cheap attention types: local, where each token attends a window of w neighbors, order n w, with a receptive field that grows with depth like convolutions so after L layers a token indirectly sees about L times w tokens, plus Longformer’s dilated windows to widen the field without more compute; global, where a small set of g positions attend to every position and every position attends to them, a full row and full column in the attention matrix, chosen by the task such as the CLS token for classification or question tokens for QA, acting as relays so any two far-apart tokens communicate in two hops token to global to token at order g n which is order n; and for BigBird random, where each token attends to r random positions, giving the sparse graph small-world connectivity so short paths connect every pair with high probability. Window plus global plus random is BigBird’s full recipe, total edges order n times w plus g plus r equals order n. BigBird’s authors proved this sparse attention is a universal approximator of sequence-to-sequence functions and is Turing complete, retaining the expressive power of full attention, turning sparse attention seems fine empirically into sparse attention provably loses nothing in principle at order n cost. The wins were concrete: 8 times longer inputs to 4096 tokens, state of the art on long-document QA and summarization, and new results on DNA modeling since a genome is a very long sequence. Caveats: global tokens are chosen not learned so picking wrong leaves a token without needed global reach; two non-global non-window tokens still cannot attend directly in one layer and route through a global token or random edge indirectly; irregular patterns fight hardware especially random attention so BigBird uses block-random and blocked windows; and FlashAttention making exact attention cheap narrowed the need at moderate lengths though window plus global remains a backbone of long-context and encoder models. In the lineage Longformer and BigBird refine the sparse family with an explicit global-relay mechanism and for BigBird a theory of why it suffices; their local half lives on in sliding-window attention in Mistral, and the global-token idea recurs as attention sinks and register tokens, while the low-rank and kernel branches reach linearity by approximating the whole matrix rather than sparsifying it. Two interactive widgets let the reader toggle and size the local window, global tokens, and random edges on a 24 by 24 attention matrix and watch the edge count stay far below full n squared with global tokens keeping everything reachable in two hops, and compare sparse window plus global plus random edges against full n squared as documents get long.