Akshath Tiwari

This is an interactive, derivation-first learning map that traces the evolution of transformer and large language model architecture as one strictly cumulative path, from a 2019 decoder-only baseline to the 2.8-trillion-parameter Kimi K3. It is organized into ten phases. Phase 0 (Bedrock) covers language modeling as next-token prediction — the chain rule of probability, the autoregressive factorization, negative log-likelihood and cross-entropy, and perplexity — plus tokenization (character vs word vs subword, the byte-pair-encoding merge algorithm, byte-level BPE in GPT-2, SentencePiece and Unigram), token embeddings (one-hot to dense lookup, embedding geometry, weight tying, the unembedding and softmax head), and the pre-transformer era (n-grams, RNNs, LSTMs, sequence-to-sequence, Bahdanau attention, and the bottlenecks that motivated Attention Is All You Need). Phase 1 (Attention) derives scaled dot-product attention including the square-root-of-d_k variance argument and a worked three-token example, causal masking and teacher forcing, multi-head attention with the output projection and tensor shapes, and the quadratic O(n squared) complexity that drives the rest of the field. Phase 2 covers the transformer block: residual stream, normalization (post-LN, pre-LN, RMSNorm), the feed-forward network and activation evolution (GELU, SwiGLU, SiTU), and the full GPT-2 anatomy. Phase 3 covers positional information: absolute and relative position encodings, rotary position embeddings (RoPE), and long-context scaling (position interpolation, NTK-aware scaling, YaRN, ALiBi) up to one million tokens. Phase 4 is attention efficiency: the KV cache, multi-query and grouped-query attention, FlashAttention, sliding-window and sparse attention, multi-head latent attention (MLA), the linear-attention and state-space lineage (Mamba, DeltaNet), and Kimi Delta Attention (KDA). Phase 5 covers Mixture of Experts: fundamentals and routing, load balancing, fine-grained and shared experts, and Kimi K3’s Stable LatentMoE. Phase 6 covers pretraining: scaling laws, data pipelines, optimizers (AdamW, Muon, MuonClip), precision (BF16, FP8, MXFP4 quantization-aware training), parallelism, and training stability. Phase 7 covers post-training and reasoning: supervised fine-tuning and LoRA, RLHF and DPO, reinforcement learning with verifiable rewards and GRPO, and agentic training. Phase 8 covers inference and serving: prefill versus decode, serving systems (continuous batching, PagedAttention, RadixAttention), speculative decoding, quantization, multi-token prediction, and serving trillion-parameter MoE models. Phase 9 is a model evolution timeline from GPT-2 and GPT-3 through Chinchilla, LLaMA, Mistral, Mixtral, DeepSeek-V2, DeepSeek-V3 and R1, Qwen3, Kimi K2, Kimi Linear, and Kimi K3. The map is progressively filled: the Bedrock and Attention phases are worked in full detail, and the remaining phases carry orientations that are deepened one phase at a time. Kimi K3 specifics are based on what is publicly known as of July 2026; the full technical report lands on July 27, 2026.