Akshath Tiwari

signal chain notes · attention mechanics

Every attention layer has a gain knob.
Most tutorials skip it.

Why transformers divide by √dk before the softmax, traced through variance, clipping, and one automatic gain control that fixes both.

↓ session starts below

track 01 · line check

Meet the session players

Before any signal reaches the mixing bus, three channels are patched in for every token in the sequence. Each one is just a learned linear projection of the same input embedding, sent down a different cable.

Channel strip

CH 1

Query

What this token is looking for. One query per position, asking the room a question.

CH 2

Key

What this token is offering to be matched against. Every position broadcasts one.

CH 3

Value

What actually gets blended into the output once the match is scored.

The score between a query and a key is their dot product, a single number that says how well one token's question matches another token's offer. Stack those scores across every key, run them through softmax, and you get a mixing ratio: how much of each Value gets pulled into the output. That mixing ratio is attention. It sounds simple, and it is, right up until you ask what happens to that dot product as the channel gets wider.

track 02 · the build-up

Stack enough channels and the meter climbs on its own

Nothing about any individual component changes. The meter still climbs.

Fresh out of layer norm, each entry of q andk is well behaved: mean 0, variance 1, roughly independent of its neighbours. The dot product just sums the products of matching entries:

q⋅k=∑i=1dkqikiq \cdot k = \sum_{i=1}^{d_k} q_i k_i

Look at one term. Since qi andki are independent with mean 0, that term has mean 0 too, so there is no systematic bias in either direction. But its variance is 1, the same as every other term, regardless of how many dimensions the vectors have:

Var(qiki)=E[qi2] E[ki2]=1×1=1\mathrm{Var}(q_i k_i) = E[q_i^2]\,E[k_i^2] = 1 \times 1 = 1

Here is the part that catches people out. When you sum independent random variables, variances add, not standard deviations. Every one of those small, well behaved terms is contributing noise that does not cancel out in aggregate:

Var(q⋅k)=∑i=1dkVar(qiki)=dk\mathrm{Var}(q \cdot k) = \sum_{i=1}^{d_k} \mathrm{Var}(q_i k_i) = d_k

Nudge the channel count fader below and watch the readout. Each square is one dimension, contributing exactly one unit of variance. Nothing about any single square changes as you add more of them. What changes is the total.

Variance accumulator

Channel count dk64
481632641282565121024
Var(q·k)
64
σ = √dk
8.00

At dk = 4 the standard deviation is a tame 2. Bydk = 512 it is over 22, and that is the raw score two tokens get compared with, before softmax has done anything at all.

track 03 · bench test

Do not take the algebra's word for it. Measure it.

The line Var(q·k) = dk is easy to write and easy to distrust. So here is the bench version. The signal generator below draws two fresh random vectors of dimension dk, each entry sampled mean 0 and variance 1, takes their dot product, and drops that single number into a histogram. Then it does it a few thousand more times.

Signal generator · empirical bench

Channel count dk4
41664256
−400+40
0 samples
measured σ
0.00
theory √dk
2.00
measured Var
0.00
theory dk
4

The axis does not move. The spread does. At dk = 4 the dot products pile up in a tight spike near zero. Step the fader up and the same axis has to hold a much wider pile, because every extra dimension throws one more unit of variance onto a sum that never learns to cancel. Measured σ tracks √dkevery time, which is the theory from Track 02 landing exactly where it said it would.

track 04 · clipping

Turn the gain up and the mix collapses to one channel

Softmax is scale sensitive. Feed it logits with a wide spread and it does not distribute attention gracefully. It slams almost all the weight onto whichever token happened to score highest and drives everyone else toward zero. In audio terms: the signal clips.

Below is a query token scoring five candidate keys. The raw scores are fixed. What changes is only the dimension the vectors live in, which sets how hard those scores get amplified before softmax sees them.

Softmax bus · no gain control

Channel count dk64
481632641282565121024
logit σ
5.84
top weight
98%
grad at top ≈ s(1−s)
0.018
Logit spread
safesaturatingclipped

Push the fader past a few hundred dimensions and one channel pins near 100% while the rest go dark. That is not the model getting more confident. It is the softmax input overflowing its useful range. And the gradient of softmax involves a term likesi(1−si): oncesi is pinned near 0 or 1, that term collapses toward zero too. The model cannot learn to redistribute attention, because there is almost no gradient left to redistribute it with.

track 05 · the fix

The fix is one automatic gain control, not a bigger desk

You do not need a wider dynamic range or a cleverer softmax. You need to undo the one thing that is actually growing: divide the score by √dkbefore it reaches softmax.

Var ⁣(q⋅kdk)=Var(q⋅k)dk=dkdk=1\mathrm{Var}\!\left(\frac{q \cdot k}{\sqrt{d_k}}\right) = \frac{\mathrm{Var}(q \cdot k)}{d_k} = \frac{d_k}{d_k} = 1

That is the whole mechanism. It is not a bigger mixing desk or a smarter softmax. It is a gain stage that exactly cancels the √dk growth, no matter how many channels you are summing over. Flip the switch below on the same rig from Track 03 and push the fader as far as it goes.

Softmax bus · with gain control

divide by √dk
Channel count dk64
481632641282565121024
logit σ
0.73
top weight
43%
grad at top ≈ s(1−s)
0.246
Logit spread
safesaturatingclipped

With the gain control engaged, the reading barely moves as you drag the fader from 4 dimensions to 1024. That is the invariance in the formula above, made visible: the correction depends only on dk itself, so it scales exactly as fast as the problem does.

track 06 · why the square root

Too little gain, just right, too much

Here is the question that always comes next: if the variance grew by a factor ofdk, why undo it with √dkinstead of the whole dk? Because you are correcting the standard deviation, not the variance. Variance grows like dk, so the spread you actually feel grows like its square root. Divide by that same square root and the spread returns to 1. Divide by the full dk and you overshoot in the other direction.

Three buses, one dimension fader, same raw scores. Watch all three at once as you pushdk up.

A/B/C · three divisors, one signal

Channel count dk256
481632641282565121024
÷ 1clips
logit σ 0.00
÷ √dkbalanced
logit σ 0.00
÷ dkwashed out
logit σ 0.00

Both extremes destroy the same thing: the model's ability to tell keys apart. On the left, one channel pins to 100% and the differences that survive are all noise. On the right, dividing by the full dk shrinks every logit toward zero, softmax flattens to near-uniform, and every key gets roughly the same weight no matter how well it matched. Only the middle bus keeps the spread near 1, where the differences the model learned are the differences softmax actually acts on. That is the whole reason the exponent is½ and not 1.

track 07 · full mixdown

One rig, every control

Everything from the last three tracks, live on one console. Switch queries, drag the dimension fader, toggle the gain control, and watch the logit spread, the gradient, and the channel bank respond together.

Master bus

divide by √dk
Channel count dk64
481632641282565121024
logit σ
0.73
top weight
43%
grad at top ≈ s(1−s)
0.246
Logit spread
safesaturatingclipped

Try both queries with the gain control off. One starts peaked, one starts nearly flat. It does not matter. Once the dimension fader climbs far enough, both collapse to a single lit channel, because the differences between raw scores get amplified by the same√dk factor as everything else. The noise does not care what the underlying signal looked like.

track 08 · patch notes

You have seen this knob before: it is temperature

Strip the analogy away and √dk is a constant that divides the logits before softmax. That is the exact shape of a softmax temperature:

softmax ⁣(zτ)withτ=dk\mathrm{softmax}\!\left(\frac{z}{\tau}\right)\quad\text{with}\quad \tau = \sqrt{d_k}

Temperature is the one dial that stretches or compresses the gap between logits without touching which one is largest. High temperature flattens the distribution toward uniform. Low temperature sharpens it toward one-hot. Scaled attention is just softmax run at a fixed temperature of √dk, chosen so the logit spread lands near 1 no matter how wide the head is. The architecture sets it once, up front, instead of leaving it to be tuned.

If you have ever reached for temperature scaling to calibrate a classifier's confidence, or dialed sampling temperature on a generation head, this is the same lever seen from the other end. Same mechanism, different point in the pipeline: one is picked by hand after training to fix over-confident probabilities, the other is baked into the forward pass to keep gradients alive during it. Once you see the √dk as a temperature, it stops looking like a magic constant in a paper and starts looking like the one setting that keeps softmax in its usable range.