Akshath Tiwari

This is a from-scratch, complete-math explainer of scaled dot-product attention, the core operation of the transformer. Its central idea is that attention is a differentiable, content-addressed dictionary lookup: each position emits a query asking what it is looking for, every position exposes a key describing what it matches and a value describing what it will hand over, the query is compared against every key by dot product to get relevance weights, and the output is the weighted average of the values, with a division by the square root of d_k to keep the lookup soft enough to train. It begins with a disambiguation example, the word bank needing to attend to river or to account depending on context, motivating content-based rather than position-based gathering. The query, key, and value projections are learned linear maps of the input, Q equals X times W_Q, K equals X times W_K, and V equals X times W_V, each producing an n by d_k or n by d_v matrix with one row per token, using three separate matrices because asking, matching, and providing are genuinely different roles; in the original transformer d equals 512 is split across eight heads so each head uses d_k equals d_v equals 64. The dot product measures similarity because it factors into the product of the vector magnitudes and the cosine of the angle between them, making it an unnormalized cosine similarity that training shapes so a token’s query aligns with the keys of the tokens it should attend to. Computing this for all pairs gives the score matrix S equals Q times K transpose, an n by n matrix whose entry i j is the relevance of token j to token i, which is both the one-step path between any two positions and the source of the quadratic cost. The scaling by the square root of d_k is derived in full: assuming query and key components are independent with mean zero and variance one, the dot product has mean zero, and its variance is the sum over d_k independent terms each with variance one, so the variance equals d_k and the standard deviation equals the square root of d_k; with d_k equals 64 the standard deviation is eight, large enough to push the softmax into a saturated regime where one weight is near one, the rest near zero, and the gradient is near zero, stalling learning, so dividing by the square root of d_k rescales the variance back to one. Row-wise softmax turns each row of the scaled scores into a probability distribution summing to one, token i’s attention distribution over all tokens, and the output is the weighted sum of value vectors, giving the full formula Attention of Q, K, V equals softmax of Q K transpose over the square root of d_k, times V. A fully worked 3-token example uses d equals d_k equals d_v equals 2 with inputs x1 equals one zero, x2 equals zero one, x3 equals one one, and specific weight matrices, computing Q, K, V, then the score matrix, then dividing by the square root of two, then row-wise softmax giving the attention matrix, then the output as the attention-weighted value average; token two draws 57.6 percent of its output from value three and ends at 0.860, 1.432, all cross-checked with a script. A causal mask for language models sets future scores to negative infinity before the softmax so each token attends only to itself and the past, the entire difference between an encoder and a decoder attention layer. Caveats: attention is quadratic in sequence length, a single head averages away detail so real models use multi-head attention, attention weights are not faithful explanations, and the variance argument is an idealization that holds at initialization. It closes by pointing to multi-head attention and the full transformer block with residual connections, layer norm, and a feed-forward network. Two interactive widgets let the reader crank d_k and watch the unscaled softmax collapse to a one-hot while the scaled version stays soft, and steer a query vector over three fixed keys to watch content-based selection and the value blend.