This is a deep explainer of residual connections in transformers and how Kimi K3 extends them. Its central idea is that a residual connection reframes a layer from computing a fresh representation to computing a small edit F of x added to a shared running representation, x becomes x plus F of x, and that one addition does three things: it makes very deep networks trainable because doing nothing means F goes to zero which is easy, it gives gradients a direct highway to every layer, and it turns the running x into the residual stream, a shared bus that every layer reads from and writes to and the central object of mechanistic interpretability. It opens with the degradation problem from He and colleagues 2015: a deeper network should never be worse than a shallower one since it could set extra layers to identity, yet plain deep nets trained worse, higher training error, because they could not even learn the identity. ResNet fixes this by learning the residual function F of x equals H of x minus x, so doing nothing means F equals zero rather than forcing a stack of nonlinear layers to reproduce their input, which the paper says makes residual networks easier to optimize, enabling 152 layers. The gradient highway is derived: with x at layer l equals x at l minus 1 plus F of x at l minus 1, the derivative of one block is I plus the Jacobian of F, and chaining across layers gives a product of I plus Jacobian terms whose expansion contains a clean identity term, so the gradient reaching an early layer always retains an undamped copy of the gradient from the top; a plain network instead has a bare product of Jacobians that vanishes exponentially with depth if their norm is below one or explodes if above one. Identity initialization means initializing F near zero so the network begins as the identity and training opts into depth gradually, each block learning a small perturbation. The residual stream view from Elhage and colleagues 2021 reframes a transformer: every sublayer is x plus sublayer of x, so a single vector threads the whole depth and every component only reads from it and writes back by addition; this is the dominant interpretability mental model because the stream is linear, modified only by addition so contributions decompose into a clean sum and features are directions, and because it is a communication channel where heads in different layers compose by leaving messages, exactly the induction-head circuit. The limitation is that the standard residual stream weights every earlier layer equally, a flat sum, so with great depth early contributions dilute into an undifferentiated blend. Kimi K3, an open 2.8 trillion parameter model from Moonshot in 2026, addresses this with Attention Residuals or AttnRes, which replace the fixed equal-weighted sum over previous layers with a softmax attention over previous layers outputs: the current layer forms a query, earlier layers representations are keys and values, and each layer builds its residual input as a learned input-dependent weighted combination of the depth history, x at l minus 1 equals attention of query l over keys and values from layers zero to l minus 1. It is the same query key value primitive as ordinary attention but pointed at the layer axis instead of the token axis, letting deep layers selectively route forward relevant earlier representations. Because attending over all previous layers is quadratic in depth, order L squared d, the same all-pairs cost on the layer axis, K3 uses Block AttnRes, partitioning layers into a handful of blocks on the order of eight and attending over block-level representations, dropping the overhead to roughly linear in depth and bounding inference state. Caveats: AttnRes trades the free single addition for parameters, compute, and a quadratic term that Block AttnRes mitigates; it complicates the clean linear residual-stream decomposition because the depth mixing becomes input-dependent and nonlinear; the reported efficiency gains vary across the launch-window sources so the mechanism is the reliable takeaway and exact percentages are provisional; and non-uniform depth connectivity has prior art in DenseNets, highway networks, and gating. It closes by tying the series together: attention is a general operation, form a query, score against keys, read a softmax-weighted average of values, that the field re-points at new axes, tokens then heads then layers, with the residual stream as the shared medium. Two interactive widgets let the reader vary depth and per-layer gain to see the plain network gradient vanish or explode while the residual identity path stays stable, and toggle standard flat residual versus AttnRes to compare how a layer weights the depth history below it.