signal chain notes · attention mechanics
Every attention layer has a gain knob.
Most tutorials skip it.
Why transformers divide by √dk before the softmax, traced through variance, clipping, and one automatic gain control that fixes both.
↓ session starts below
track 01 · line check
Meet the session players
Before any signal reaches the mixing bus, three channels are patched in for every token in the sequence. Each one is just a learned linear projection of the same input embedding, sent down a different cable.
Channel strip
Query
What this token is looking for. One query per position, asking the room a question.
Key
What this token is offering to be matched against. Every position broadcasts one.
Value
What actually gets blended into the output once the match is scored.
The score between a query and a key is their dot product, a single number that says how well one token's question matches another token's offer. Stack those scores across every key, run them through softmax, and you get a mixing ratio: how much of each Value gets pulled into the output. That mixing ratio is attention. It sounds simple, and it is, right up until you ask what happens to that dot product as the channel gets wider.
track 02 · the build-up
Stack enough channels and the meter climbs on its own
Nothing about any individual component changes. The meter still climbs.
Fresh out of layer norm, each entry of q andk is well behaved: mean 0, variance 1, roughly independent of its neighbours. The dot product just sums the products of matching entries:
Look at one term. Since qi andki are independent with mean 0, that term has mean 0 too, so there is no systematic bias in either direction. But its variance is 1, the same as every other term, regardless of how many dimensions the vectors have:
Here is the part that catches people out. When you sum independent random variables, variances add, not standard deviations. Every one of those small, well behaved terms is contributing noise that does not cancel out in aggregate:
Nudge the channel count fader below and watch the readout. Each square is one dimension, contributing exactly one unit of variance. Nothing about any single square changes as you add more of them. What changes is the total.
Variance accumulator
At dk = 4 the standard deviation is a tame 2. Bydk = 512 it is over 22, and that is the raw score two tokens get compared with, before softmax has done anything at all.
track 03 · bench test
Do not take the algebra's word for it. Measure it.
The line Var(q·k) = dk is easy to write and easy to distrust. So here is the bench version. The signal generator below draws two fresh random vectors of dimension dk, each entry sampled mean 0 and variance 1, takes their dot product, and drops that single number into a histogram. Then it does it a few thousand more times.
Signal generator · empirical bench
The axis does not move. The spread does. At dk = 4 the dot products pile up in a tight spike near zero. Step the fader up and the same axis has to hold a much wider pile, because every extra dimension throws one more unit of variance onto a sum that never learns to cancel. Measured σ tracks √dkevery time, which is the theory from Track 02 landing exactly where it said it would.
track 04 · clipping
Turn the gain up and the mix collapses to one channel
Softmax is scale sensitive. Feed it logits with a wide spread and it does not distribute attention gracefully. It slams almost all the weight onto whichever token happened to score highest and drives everyone else toward zero. In audio terms: the signal clips.
Below is a query token scoring five candidate keys. The raw scores are fixed. What changes is only the dimension the vectors live in, which sets how hard those scores get amplified before softmax sees them.
Softmax bus · no gain control
Push the fader past a few hundred dimensions and one channel pins near 100% while the rest go dark. That is not the model getting more confident. It is the softmax input overflowing its useful range. And the gradient of softmax involves a term likesi(1−si): oncesi is pinned near 0 or 1, that term collapses toward zero too. The model cannot learn to redistribute attention, because there is almost no gradient left to redistribute it with.
track 05 · the fix
The fix is one automatic gain control, not a bigger desk
You do not need a wider dynamic range or a cleverer softmax. You need to undo the one thing that is actually growing: divide the score by √dkbefore it reaches softmax.
That is the whole mechanism. It is not a bigger mixing desk or a smarter softmax. It is a gain stage that exactly cancels the √dk growth, no matter how many channels you are summing over. Flip the switch below on the same rig from Track 03 and push the fader as far as it goes.
Softmax bus · with gain control
With the gain control engaged, the reading barely moves as you drag the fader from 4 dimensions to 1024. That is the invariance in the formula above, made visible: the correction depends only on dk itself, so it scales exactly as fast as the problem does.
track 06 · why the square root
Too little gain, just right, too much
Here is the question that always comes next: if the variance grew by a factor ofdk, why undo it with √dkinstead of the whole dk? Because you are correcting the standard deviation, not the variance. Variance grows like dk, so the spread you actually feel grows like its square root. Divide by that same square root and the spread returns to 1. Divide by the full dk and you overshoot in the other direction.
Three buses, one dimension fader, same raw scores. Watch all three at once as you pushdk up.
A/B/C · three divisors, one signal
Both extremes destroy the same thing: the model's ability to tell keys apart. On the left, one channel pins to 100% and the differences that survive are all noise. On the right, dividing by the full dk shrinks every logit toward zero, softmax flattens to near-uniform, and every key gets roughly the same weight no matter how well it matched. Only the middle bus keeps the spread near 1, where the differences the model learned are the differences softmax actually acts on. That is the whole reason the exponent is½ and not 1.
track 07 · full mixdown
One rig, every control
Everything from the last three tracks, live on one console. Switch queries, drag the dimension fader, toggle the gain control, and watch the logit spread, the gradient, and the channel bank respond together.
Master bus
Try both queries with the gain control off. One starts peaked, one starts nearly flat. It does not matter. Once the dimension fader climbs far enough, both collapse to a single lit channel, because the differences between raw scores get amplified by the same√dk factor as everything else. The noise does not care what the underlying signal looked like.
track 08 · patch notes
You have seen this knob before: it is temperature
Strip the analogy away and √dk is a constant that divides the logits before softmax. That is the exact shape of a softmax temperature:
Temperature is the one dial that stretches or compresses the gap between logits without touching which one is largest. High temperature flattens the distribution toward uniform. Low temperature sharpens it toward one-hot. Scaled attention is just softmax run at a fixed temperature of √dk, chosen so the logit spread lands near 1 no matter how wide the head is. The architecture sets it once, up front, instead of leaving it to be tuned.
If you have ever reached for temperature scaling to calibrate a classifier's confidence, or dialed sampling temperature on a generation head, this is the same lever seen from the other end. Same mechanism, different point in the pipeline: one is picked by hand after training to fix over-confident probabilities, the other is baked into the forward pass to keep gradients alive during it. Once you see the √dk as a temperature, it stops looking like a magic constant in a paper and starts looking like the one setting that keeps softmax in its usable range.