Blog
Notes on fine-tuning, agents, evaluation, and everything in between.
September 5, 2026
Compression Is Intelligence: Why Next-Token Prediction Is Secretly Data Compression
A ground-up derivation of self-information, entropy, and cross-entropy — and why an LLM's cross-entropy pre-training loss is literally the number of bits it takes to compress its data. From a moon rover's instruction code to DeepMind's Chinchilla out-compressing PNG and FLAC. Interactive, inspired by 3Blue1Brown's Reinventing Entropy.
July 29, 2026
GEPA: Reading the Film Instead of the Scoreboard
A deep dive on GEPA (Genetic-Pareto), the reflective prompt optimizer that can beat reinforcement learning while using up to 35x fewer rollouts. Instead of learning from a single scalar reward, GEPA hands the whole execution trace to an LLM that diagnoses failures in natural language and rewrites the prompt, and it keeps a Pareto roster of prompts where each is the best at some instance rather than chasing a single best-on-average winner. With an interactive greedy-vs-Pareto selection explorer and a sample-efficiency comparator.
July 29, 2026
SkillOpt: A Learning Rate for Your Prompt
A deep dive on SkillOpt (Microsoft Research, arXiv 2605.23904), which treats prompt optimization as gradient descent over text. The skill document is the parameter, the task score is the loss, an optimizer LLM's add/delete/replace edits are the textual gradient, and the number of edits allowed per step is the learning rate. With a strict held-out validation gate for acceptance, a rejected-edit buffer for momentum, and a cosine edit-budget schedule, it lifts a GPT-5.5 agent from 41.8 to 80.7 percent on SpreadsheetBench by editing a single file. Includes an interactive learning-rate-overshoot demo and a step-through of one optimization step. Companion to the GEPA deep dive.
July 28, 2026
The Quadratic Wall: Why Attention is O(n²), Where It Bites, and the Decade of Research It Drove
A rigorous derivation of attention's O(n²·d) compute and O(n²) memory from the matrix shapes, the crucial compute-versus-memory distinction (and where FlashAttention fits), where the quadratic bites in compute-bound prefill versus memory-bound decode, and how this single fact drove nearly all attention-architecture research from Sparse Transformers in 2019 to Kimi K3 — with an interactive cost explorer and a filterable efficiency-lineage map.
July 28, 2026
FlashAttention: Attention is Memory-Bound, and the Online-Softmax Fix
Why attention's wall-clock cost is memory traffic, not FLOPs — the GPU SRAM-vs-HBM hierarchy, how FlashAttention tiles Q/K/V into fast on-chip memory and never materializes the n×n matrix, the online-softmax recurrence (running max and running sum) that makes tiling and softmax compatible while staying exact, backward-pass recomputation, and the 2-4x exact speedup that made it the default — with an interactive HBM-traffic explorer and a streaming-softmax stepper.
July 28, 2026
Kimi Delta Attention: The Delta Rule, an Editable State, and the Frontier Hybrid
The finale of the attention-efficiency series: how Kimi Delta Attention fixes linear attention's recall weakness by making the fixed-size state editable with the delta rule (read the old value, subtract it, write the correction), adds per-channel gated forgetting on top of Gated DeltaNet, and hybridizes 3 KDA linear layers to 1 full-attention MLA layer — the engine of Moonshot's Kimi Linear and the 2.8-trillion-parameter Kimi K3, which synthesizes nearly every idea in the series. With an interactive delta-rule associative-memory demo and a hybrid-stack visualizer.
July 28, 2026
Linformer: Self-Attention is Low-Rank, and What That Buys You (O(n))
How the Linformer reaches linear-complexity attention by exploiting a measured fact — the n×n self-attention matrix is approximately low-rank. The spectral evidence and Johnson-Lindenstrauss argument, the two learned projection matrices E and F that compress keys and values from length n to k, why the n×k attention costs O(n), k=128-256 matching RoBERTa, and the fixed-length limitation — with an interactive low-rank reconstruction demo and a cost/shape explorer.
July 28, 2026
Longformer & BigBird: Local Windows, Global Tokens, and Random Edges
How Longformer and BigBird scale transformers to long documents with linear attention — sliding-window local attention, a few global tokens that attend to and from everything as two-hop relays, and BigBird's random edges plus the proof that window+global+random is a universal approximator of full attention — with an interactive attention-pattern builder and a cost explorer.
July 28, 2026
Mamba & State-Space Models: Trading Perfect Recall for a Fixed-Size State
How state-space models escape attention's O(n²) by compressing all history into a fixed-size state (O(1) memory, O(n) compute) — the linear SSM recurrence, why it trains as a convolution but runs as a recurrence, why fixed (time-invariant) dynamics failed on language, and how Mamba's input-dependent selection recovers content-based memory with a hardware-aware parallel scan — plus the recall-vs-efficiency tradeoff, hybrids, and the SSM-linear-attention duality. With an attention-vs-SSM tradeoff view and an interactive selective-state stepper.
July 28, 2026
Multi-head Latent Attention: Compress the KV Cache Instead of Sharing It
How DeepSeek's Multi-head Latent Attention shrinks the KV cache by low-rank compression rather than head-sharing — down-project each token to a small latent, cache only that, and up-project it back to full per-head keys and values so head diversity is preserved and quality matches MHA. The mechanism, the decoupled-RoPE fix for the incompatibility with rotary embeddings, and why MLA's cache is smaller than GQA while quality stays at MHA level (DeepSeek-V2/V3). With an interactive MHA-vs-GQA-vs-MLA cache comparator and a compress-cache-reconstruct pipeline.
July 28, 2026
Multi-Query & Grouped-Query Attention: Shrinking the KV Cache
Why autoregressive decoding is bottlenecked by the KV cache rather than attention FLOPs, and how Multi-Query Attention (one shared key/value head) and Grouped-Query Attention (one K/V head per group) shrink that cache by a factor of h or h/G — the mechanism, the mean-pool uptraining recipe, adoption in Llama 2 / Mistral, and the quality-vs-cache trade — with an interactive head-grouping KV-cache calculator and a cache-vs-context explorer.
July 28, 2026
Multi-Head Attention in Full: Head Splitting, W_O, What Heads Learn, and Exact Tensor Shapes
A complete account of multi-head attention — why multiple heads beat one big head, head-dimension splitting (d_k = d_model / h), the concatenation and output projection W_O, what individual heads learn (previous-token, positional, and induction heads), and every tensor shape from input to output — with an interactive shape-and-parameter calculator and a per-head attention viewer.
July 28, 2026
Performer: Linearizing Softmax Attention with Kernel Feature Maps (FAVOR+)
How the Performer reaches O(n) attention by never forming the n×n matrix. Why the softmax blocks matrix-multiply associativity, how exp(q·k) as a kernel factorizes into feature maps so you can compute phi(K)^T V once and let every query read it, and how FAVOR+ (Fast Attention Via positive Orthogonal Random features) builds an unbiased estimator of the softmax kernel — with an interactive associativity-reorder cost comparator and a live FAVOR+ approximation demo.
July 28, 2026
Reformer: Attention as Nearest-Neighbor Search, LSH Bucketing, and the O(n log n) Trick
How the Reformer turns attention into an approximate nearest-neighbor search — the softmax-concentration insight, angular locality-sensitive hashing that buckets similar queries and keys, why bucketed attention is O(n log n), the shared-QK and multi-round hashing details, and reversible layers for memory — with an interactive angular-LSH bucketing ring.
July 28, 2026
The Residual Stream: Skip Connections, Gradient Highways, Interpretability, and Kimi K3's Attention Residuals
Residual connections deeply — the residual stream view, why gradients flow through skip connections (derived), the identity-initialization intuition, and how the residual stream became the dominant mental model in mechanistic interpretability — then how Kimi K3's Attention Residuals (AttnRes) apply attention across depth instead of a flat residual sum, with an interactive gradient-flow demo and a depth-attention comparison.
July 28, 2026
Sliding Window Attention: Depth Buys Range, and a KV Cache That Never Grows
How Mistral's sliding window attention gets effectively-long context from a small fixed window — each token attends only to the last W tokens (O(n·W) compute), but stacking L layers gives an L×W receptive field because information hops one window per layer like a CNN, and because no token looks past W the KV cache becomes a fixed-size rolling buffer that never grows. With an interactive receptive-field-vs-depth visualizer and a rolling-buffer cache demo.
July 28, 2026
The Sparse Transformer: Strided and Fixed Attention, Two-Hop Reachability, and the O(n√n) Escape
How the 2019 Sparse Transformer became the first serious escape from attention's quadratic cost — factorized strided and fixed sparse attention patterns, the two-hop rook-style reachability that keeps the sequence globally connected, the exact O(n√n) derivation from minimizing the per-token cost at stride √n, and its place at the head of the efficient-attention lineage — with an interactive pattern-grid visualizer and a √n cost explorer.
July 27, 2026
Causal Masking in Decoder-Only Transformers: the −∞ Mask, Teacher Forcing, and Training vs Inference
Why autoregressive generation needs a causal mask, how the −∞ mask is applied before the softmax so future tokens get exactly zero attention weight, how one masked forward pass trains every position in parallel via teacher forcing, how training differs from token-by-token inference, and where exposure bias comes from — with an interactive masked-attention matrix and a training-vs-inference stepper.
July 27, 2026
Scaled Dot-Product Attention from Scratch: Q, K, V, the √dₖ Derivation, and a Worked Example
A complete, from-scratch derivation of scaled dot-product attention — what the query, key, and value projections actually are, why the dot product measures similarity, why we divide by the square root of d_k (with the full variance argument), row-wise softmax, and a fully worked 3-token numeric example — plus interactive demos for the √dₖ saturation effect and content-based query selection.
July 26, 2026
Language Modeling from Scratch: Next-Token Prediction, Cross-Entropy, and Perplexity
A from-scratch derivation of how language models work as next-token predictors — the chain rule of probability, the autoregressive factorization, cross-entropy and negative log-likelihood as one quantity, and perplexity as an effective branching factor — with a fully worked example where every number cross-checks and an interactive surprise ledger.
July 26, 2026
Before the Transformer: N-grams, RNNs, LSTMs, Seq2Seq, and the Bottlenecks That Forced Attention
The pre-transformer lineage told as one problem attacked four times — carrying information across a long sequence. N-gram Markov models, RNNs and the vanishing gradient, LSTMs and the gated cell state, seq2seq and the fixed-vector bottleneck, Bahdanau attention, and exactly which bottlenecks (sequential computation and path length) motivated Attention Is All You Need in 2017 — with an interactive path-length comparator and alignment demo.
July 26, 2026
Token Embeddings from Scratch: One-Hot to Dense Vectors, Weight Tying, and the Softmax Head
How a language model turns integer token IDs into meaning — one-hot vectors and why they fail, the embedding matrix as a lookup table, what the embedding dimensions mean geometrically, tied input/output embeddings (weight tying), and the unembedding plus softmax head — with a worked example and an interactive 2-D embedding map you can steer.
July 26, 2026
Tokenization from Scratch: BPE, Byte-Level BPE, Unigram, and Why LLMs Fail at Spelling
A full-depth walk through tokenization — character vs word vs subword, the Byte-Pair Encoding algorithm step by step with a worked merge example you can run, byte-level BPE (GPT-2's 50,257 vocabulary), WordPiece and Unigram/SentencePiece, the vocabulary-size trade-off, and why tokenization is behind LLM failures at spelling, arithmetic, glitch tokens, and multilingual cost.
July 25, 2026
Transformer Evolution: A Derivation-First Map from GPT-2 to Kimi K3
An interactive, strictly cumulative learning map for transformer and LLM internals — from next-token prediction and scaled dot-product attention through KV caching, MoE, pretraining, post-training and serving, ending at the Kimi K3 teardown. Tap any card to expand its full worked derivation.
July 20, 2026
DFlash: Draft a Whole Block in One Shot
Block-diffusion drafting for speculative decoding. DFlash replaces the sequential drafter with a diffusion adapter that fills a whole block of tokens in one parallel pass, conditioned on the target model's own hidden features injected into every draft layer, for over 6x lossless speedup. Interactive breakdown of both levers.
July 20, 2026
DSpark: Keep the Speed, Fix What Parallel Broke
Confidence-scheduled speculative decoding with semi-autoregressive generation. DSpark stitches block coherence back with a tiny sequential head and verifies only what's worth verifying with a hardware-aware scheduler, staying lossless while shifting the serving Pareto frontier. 60 to 85% faster per user in production.
July 19, 2026
Speculative Decoding: Guess Ahead, Verify in Bulk
An interactive walk through speculative decoding: how a fast draft model and a careful verifier produce several tokens per forward pass while keeping the big model's exact output distribution. Live panels for the memory-bound bottleneck, the accept/reject rule, and the speedup ledger.
July 18, 2026
Why We Scale Attention by √dₖ
An interactive walk through the one line every attention layer runs and most tutorials skip: dividing the dot product by √dₖ before softmax. Traced through variance, softmax saturation, and vanishing gradients.
July 15, 2026
Getting a Fine-Tuned 4B Model Within 2 Points of GPT-4.1
Notes on what actually moved the needle when distilling a GPT-4.1 scoring pipeline into a fine-tuned Qwen3 model — and where the small model still falls short.
July 10, 2026
Why Completions-Only Loss Matters When Fine-Tuning Small LLMs
A practical look at masking the prompt out of the loss when fine-tuning small models for structured tasks like scoring or extraction.