This is a from-scratch explainer of token embeddings in language models. Its central idea is that a token is an arbitrary integer ID and that a token embedding is simply a row of a learned matrix selected by that ID, but because training can place every row anywhere in a d-dimensional space, the geometry of where the rows land becomes the model’s entire notion of what tokens mean. It starts with one-hot vectors as the honest first representation: a vector of length V that is one at the token’s position and zero elsewhere. One-hot vectors fail for two reasons: they are enormous and sparse (length V, which is 50,257 for GPT-2), and every pair of distinct one-hot vectors is orthogonal and equidistant, so the representation encodes zero similarity structure and the model can generalize nothing. The key reveal is that multiplying a one-hot vector by an embedding matrix E, which is V by d with one row per token, simply selects that token’s row, because every term in the sum is multiplied by zero except the hot one. So the embedding layer is an ordinary linear layer with no bias whose one-hot input degenerates the matrix multiply into a row lookup, implemented by indexing so it costs almost nothing. The vector goes from length V (sparse) to length d (dense, 768 for GPT-2 small), and the rows are free trained parameters. Geometrically, each embedding is coordinates on a learned map and the d dimensions are its axes; distance encodes similarity, so tokens appearing in similar contexts move to nearby coordinates and the model generalizes across them, and directions encode attributes, giving the word2vec analogy king minus man plus woman is approximately queen as a parallelogram of consistent offsets, with the honest caveat that this is approximate and strongest for static word2vec and GloVe embeddings rather than transformer input embeddings. The output side is the unembedding and softmax head: the transformer produces a context vector h in the same d-dimensional space, and the logit for each token v is the dot product of h with that token’s output embedding u_v, forming z equals U h, then softmax turns the logits into next-token probabilities. So prediction is a nearest-direction search: predict the token whose output embedding is most aligned with h. Weight tying sets the output matrix U equal to the input matrix E, so the logit is h dot e_v; introduced by Press and Wolf in 2016, tying halves the embedding parameters and tends to lower perplexity because the input identity and output target of a token should be related. For GPT-2 small the embedding table is V times d equals 50,257 times 768 which is about 38.6 million parameters, roughly 31 percent of the model’s 124 million total, so tying saves a second matrix of that size; tying was standard in GPT-2 and BERT, though large or multilingual models sometimes untie for flexibility, and tied matrices are biased toward the output geometry. A worked example uses d equals 2 with embeddings for cat at (2,1), dog at (1.5,1.5), run at (-1,1), and a context vector h at (1.8,1.0), giving logits 4.6, 4.2, and -0.8, and softmax probabilities 0.597, 0.400, and 0.003, so the model predicts cat then dog. Caveats: the embedding table dominates small models and scales as V times d, input embeddings are context-free starting points that attention later disambiguates, position must be added separately, and tied embeddings suffer mild representation degeneration or anisotropy. An interactive 2-D embedding map lets the reader steer the context vector h with sliders and watch the logits and softmax over six tokens, with a weight-tying toggle that separates the input map from the output detectors.