Akshath Tiwari

Blog

Notes on fine-tuning, agents, evaluation, and everything in between.

September 5, 2026

Compression Is Intelligence: Why Next-Token Prediction Is Secretly Data Compression

A ground-up derivation of self-information, entropy, and cross-entropy — and why an LLM's cross-entropy pre-training loss is literally the number of bits it takes to compress its data. From a moon rover's instruction code to DeepMind's Chinchilla out-compressing PNG and FLAC. Interactive, inspired by 3Blue1Brown's Reinventing Entropy.

information-theoryentropycross-entropycompressionllminteractive

July 29, 2026

GEPA: Reading the Film Instead of the Scoreboard

A deep dive on GEPA (Genetic-Pareto), the reflective prompt optimizer that can beat reinforcement learning while using up to 35x fewer rollouts. Instead of learning from a single scalar reward, GEPA hands the whole execution trace to an LLM that diagnoses failures in natural language and rewrites the prompt, and it keeps a Pareto roster of prompts where each is the best at some instance rather than chasing a single best-on-average winner. With an interactive greedy-vs-Pareto selection explorer and a sample-efficiency comparator.

llmprompt-optimizationreflectionreinforcement-learningdspygepaevolutionary-searchinteractive

July 29, 2026

SkillOpt: A Learning Rate for Your Prompt

A deep dive on SkillOpt (Microsoft Research, arXiv 2605.23904), which treats prompt optimization as gradient descent over text. The skill document is the parameter, the task score is the loss, an optimizer LLM's add/delete/replace edits are the textual gradient, and the number of edits allowed per step is the learning rate. With a strict held-out validation gate for acceptance, a rejected-edit buffer for momentum, and a cosine edit-budget schedule, it lifts a GPT-5.5 agent from 41.8 to 80.7 percent on SpreadsheetBench by editing a single file. Includes an interactive learning-rate-overshoot demo and a step-through of one optimization step. Companion to the GEPA deep dive.

llmprompt-optimizationgradient-descenttextual-gradientsagentsskilloptmicrosoftinteractive

July 28, 2026

The Quadratic Wall: Why Attention is O(n²), Where It Bites, and the Decade of Research It Drove

A rigorous derivation of attention's O(n²·d) compute and O(n²) memory from the matrix shapes, the crucial compute-versus-memory distinction (and where FlashAttention fits), where the quadratic bites in compute-bound prefill versus memory-bound decode, and how this single fact drove nearly all attention-architecture research from Sparse Transformers in 2019 to Kimi K3 — with an interactive cost explorer and a filterable efficiency-lineage map.

llmtransformersattentionefficiencyflashattentioninteractive

July 28, 2026

FlashAttention: Attention is Memory-Bound, and the Online-Softmax Fix

Why attention's wall-clock cost is memory traffic, not FLOPs — the GPU SRAM-vs-HBM hierarchy, how FlashAttention tiles Q/K/V into fast on-chip memory and never materializes the n×n matrix, the online-softmax recurrence (running max and running sum) that makes tiling and softmax compatible while staying exact, backward-pass recomputation, and the 2-4x exact speedup that made it the default — with an interactive HBM-traffic explorer and a streaming-softmax stepper.

llmtransformersattentionflashattentiongpuefficiencyinteractive

July 28, 2026

Kimi Delta Attention: The Delta Rule, an Editable State, and the Frontier Hybrid

The finale of the attention-efficiency series: how Kimi Delta Attention fixes linear attention's recall weakness by making the fixed-size state editable with the delta rule (read the old value, subtract it, write the correction), adds per-channel gated forgetting on top of Gated DeltaNet, and hybridizes 3 KDA linear layers to 1 full-attention MLA layer — the engine of Moonshot's Kimi Linear and the 2.8-trillion-parameter Kimi K3, which synthesizes nearly every idea in the series. With an interactive delta-rule associative-memory demo and a hybrid-stack visualizer.

llmtransformersattentionkimi-k3linear-attentiondelta-ruleefficiencyinteractive

July 28, 2026

Linformer: Self-Attention is Low-Rank, and What That Buys You (O(n))

How the Linformer reaches linear-complexity attention by exploiting a measured fact — the n×n self-attention matrix is approximately low-rank. The spectral evidence and Johnson-Lindenstrauss argument, the two learned projection matrices E and F that compress keys and values from length n to k, why the n×k attention costs O(n), k=128-256 matching RoBERTa, and the fixed-length limitation — with an interactive low-rank reconstruction demo and a cost/shape explorer.

llmtransformersattentionlinformerlow-rankefficiencyinteractive

July 28, 2026

Longformer & BigBird: Local Windows, Global Tokens, and Random Edges

How Longformer and BigBird scale transformers to long documents with linear attention — sliding-window local attention, a few global tokens that attend to and from everything as two-hop relays, and BigBird's random edges plus the proof that window+global+random is a universal approximator of full attention — with an interactive attention-pattern builder and a cost explorer.

llmtransformersattentionlongformerbigbirdsparse-attentionefficiencyinteractive

July 28, 2026

Mamba & State-Space Models: Trading Perfect Recall for a Fixed-Size State

How state-space models escape attention's O(n²) by compressing all history into a fixed-size state (O(1) memory, O(n) compute) — the linear SSM recurrence, why it trains as a convolution but runs as a recurrence, why fixed (time-invariant) dynamics failed on language, and how Mamba's input-dependent selection recovers content-based memory with a hardware-aware parallel scan — plus the recall-vs-efficiency tradeoff, hybrids, and the SSM-linear-attention duality. With an attention-vs-SSM tradeoff view and an interactive selective-state stepper.

llmtransformersattentionmambastate-space-modelsefficiencyinteractive

July 28, 2026

Multi-head Latent Attention: Compress the KV Cache Instead of Sharing It

How DeepSeek's Multi-head Latent Attention shrinks the KV cache by low-rank compression rather than head-sharing — down-project each token to a small latent, cache only that, and up-project it back to full per-head keys and values so head diversity is preserved and quality matches MHA. The mechanism, the decoupled-RoPE fix for the incompatibility with rotary embeddings, and why MLA's cache is smaller than GQA while quality stays at MHA level (DeepSeek-V2/V3). With an interactive MHA-vs-GQA-vs-MLA cache comparator and a compress-cache-reconstruct pipeline.

llmtransformersattentionmladeepseekkv-cacheefficiencyinteractive

July 28, 2026

Multi-Query & Grouped-Query Attention: Shrinking the KV Cache

Why autoregressive decoding is bottlenecked by the KV cache rather than attention FLOPs, and how Multi-Query Attention (one shared key/value head) and Grouped-Query Attention (one K/V head per group) shrink that cache by a factor of h or h/G — the mechanism, the mean-pool uptraining recipe, adoption in Llama 2 / Mistral, and the quality-vs-cache trade — with an interactive head-grouping KV-cache calculator and a cache-vs-context explorer.

llmtransformersattentionkv-cachegqamqainferenceinteractive

July 28, 2026

Multi-Head Attention in Full: Head Splitting, W_O, What Heads Learn, and Exact Tensor Shapes

A complete account of multi-head attention — why multiple heads beat one big head, head-dimension splitting (d_k = d_model / h), the concatenation and output projection W_O, what individual heads learn (previous-token, positional, and induction heads), and every tensor shape from input to output — with an interactive shape-and-parameter calculator and a per-head attention viewer.

llmtransformersattentionmulti-head-attentioninterpretabilityinteractive

July 28, 2026

Performer: Linearizing Softmax Attention with Kernel Feature Maps (FAVOR+)

How the Performer reaches O(n) attention by never forming the n×n matrix. Why the softmax blocks matrix-multiply associativity, how exp(q·k) as a kernel factorizes into feature maps so you can compute phi(K)^T V once and let every query read it, and how FAVOR+ (Fast Attention Via positive Orthogonal Random features) builds an unbiased estimator of the softmax kernel — with an interactive associativity-reorder cost comparator and a live FAVOR+ approximation demo.

llmtransformersattentionperformerlinear-attentionefficiencyinteractive

July 28, 2026

Reformer: Attention as Nearest-Neighbor Search, LSH Bucketing, and the O(n log n) Trick

How the Reformer turns attention into an approximate nearest-neighbor search — the softmax-concentration insight, angular locality-sensitive hashing that buckets similar queries and keys, why bucketed attention is O(n log n), the shared-QK and multi-round hashing details, and reversible layers for memory — with an interactive angular-LSH bucketing ring.

llmtransformersattentionreformerlshefficiencyinteractive

July 28, 2026

The Residual Stream: Skip Connections, Gradient Highways, Interpretability, and Kimi K3's Attention Residuals

Residual connections deeply — the residual stream view, why gradients flow through skip connections (derived), the identity-initialization intuition, and how the residual stream became the dominant mental model in mechanistic interpretability — then how Kimi K3's Attention Residuals (AttnRes) apply attention across depth instead of a flat residual sum, with an interactive gradient-flow demo and a depth-attention comparison.

llmtransformersresidual-connectionsinterpretabilitykimi-k3interactive

July 28, 2026

Sliding Window Attention: Depth Buys Range, and a KV Cache That Never Grows

How Mistral's sliding window attention gets effectively-long context from a small fixed window — each token attends only to the last W tokens (O(n·W) compute), but stacking L layers gives an L×W receptive field because information hops one window per layer like a CNN, and because no token looks past W the KV cache becomes a fixed-size rolling buffer that never grows. With an interactive receptive-field-vs-depth visualizer and a rolling-buffer cache demo.

llmtransformersattentionsliding-windowmistralkv-cacheefficiencyinteractive

July 28, 2026

The Sparse Transformer: Strided and Fixed Attention, Two-Hop Reachability, and the O(n√n) Escape

How the 2019 Sparse Transformer became the first serious escape from attention's quadratic cost — factorized strided and fixed sparse attention patterns, the two-hop rook-style reachability that keeps the sequence globally connected, the exact O(n√n) derivation from minimizing the per-token cost at stride √n, and its place at the head of the efficient-attention lineage — with an interactive pattern-grid visualizer and a √n cost explorer.

llmtransformersattentionsparse-attentionefficiencyinteractive

July 27, 2026

Causal Masking in Decoder-Only Transformers: the −∞ Mask, Teacher Forcing, and Training vs Inference

Why autoregressive generation needs a causal mask, how the −∞ mask is applied before the softmax so future tokens get exactly zero attention weight, how one masked forward pass trains every position in parallel via teacher forcing, how training differs from token-by-token inference, and where exposure bias comes from — with an interactive masked-attention matrix and a training-vs-inference stepper.

llmtransformersattentioncausal-maskingteacher-forcinginteractive

July 27, 2026

Scaled Dot-Product Attention from Scratch: Q, K, V, the √dₖ Derivation, and a Worked Example

A complete, from-scratch derivation of scaled dot-product attention — what the query, key, and value projections actually are, why the dot product measures similarity, why we divide by the square root of d_k (with the full variance argument), row-wise softmax, and a fully worked 3-token numeric example — plus interactive demos for the √dₖ saturation effect and content-based query selection.

llmtransformersattentionself-attentionsoftmaxinteractive

July 26, 2026

Language Modeling from Scratch: Next-Token Prediction, Cross-Entropy, and Perplexity

A from-scratch derivation of how language models work as next-token predictors — the chain rule of probability, the autoregressive factorization, cross-entropy and negative log-likelihood as one quantity, and perplexity as an effective branching factor — with a fully worked example where every number cross-checks and an interactive surprise ledger.

llmlanguage-modelingcross-entropyperplexityinformation-theoryinteractive

July 26, 2026

Before the Transformer: N-grams, RNNs, LSTMs, Seq2Seq, and the Bottlenecks That Forced Attention

The pre-transformer lineage told as one problem attacked four times — carrying information across a long sequence. N-gram Markov models, RNNs and the vanishing gradient, LSTMs and the gated cell state, seq2seq and the fixed-vector bottleneck, Bahdanau attention, and exactly which bottlenecks (sequential computation and path length) motivated Attention Is All You Need in 2017 — with an interactive path-length comparator and alignment demo.

llmtransformersrnnlstmattentionhistoryinteractive

July 26, 2026

Token Embeddings from Scratch: One-Hot to Dense Vectors, Weight Tying, and the Softmax Head

How a language model turns integer token IDs into meaning — one-hot vectors and why they fail, the embedding matrix as a lookup table, what the embedding dimensions mean geometrically, tied input/output embeddings (weight tying), and the unembedding plus softmax head — with a worked example and an interactive 2-D embedding map you can steer.

llmembeddingsweight-tyingsoftmaxrepresentationinteractive

July 26, 2026

Tokenization from Scratch: BPE, Byte-Level BPE, Unigram, and Why LLMs Fail at Spelling

A full-depth walk through tokenization — character vs word vs subword, the Byte-Pair Encoding algorithm step by step with a worked merge example you can run, byte-level BPE (GPT-2's 50,257 vocabulary), WordPiece and Unigram/SentencePiece, the vocabulary-size trade-off, and why tokenization is behind LLM failures at spelling, arithmetic, glitch tokens, and multilingual cost.

llmtokenizationbpebyte-pair-encodingsentencepieceinteractive

July 25, 2026

Transformer Evolution: A Derivation-First Map from GPT-2 to Kimi K3

An interactive, strictly cumulative learning map for transformer and LLM internals — from next-token prediction and scaled dot-product attention through KV caching, MoE, pretraining, post-training and serving, ending at the Kimi K3 teardown. Tap any card to expand its full worked derivation.

llmtransformersattentionmoelearning-mapinteractive

July 20, 2026

DFlash: Draft a Whole Block in One Shot

Block-diffusion drafting for speculative decoding. DFlash replaces the sequential drafter with a diffusion adapter that fills a whole block of tokens in one parallel pass, conditioned on the target model's own hidden features injected into every draft layer, for over 6x lossless speedup. Interactive breakdown of both levers.

llminferencespeculative-decodingdiffusioninteractive

July 20, 2026

DSpark: Keep the Speed, Fix What Parallel Broke

Confidence-scheduled speculative decoding with semi-autoregressive generation. DSpark stitches block coherence back with a tiny sequential head and verifies only what's worth verifying with a hardware-aware scheduler, staying lossless while shifting the serving Pareto frontier. 60 to 85% faster per user in production.

llminferencespeculative-decodingservinginteractive

July 19, 2026

Speculative Decoding: Guess Ahead, Verify in Bulk

An interactive walk through speculative decoding: how a fast draft model and a careful verifier produce several tokens per forward pass while keeping the big model's exact output distribution. Live panels for the memory-bound bottleneck, the accept/reject rule, and the speedup ledger.

llminferencespeculative-decodinginteractive

July 18, 2026

Why We Scale Attention by √dₖ

An interactive walk through the one line every attention layer runs and most tutorials skip: dividing the dot product by √dₖ before softmax. Traced through variance, softmax saturation, and vanishing gradients.

transformersattentioninteractivemath

July 15, 2026

Getting a Fine-Tuned 4B Model Within 2 Points of GPT-4.1

Notes on what actually moved the needle when distilling a GPT-4.1 scoring pipeline into a fine-tuned Qwen3 model — and where the small model still falls short.

llmdistillationevaluationvllm

July 10, 2026

Why Completions-Only Loss Matters When Fine-Tuning Small LLMs

A practical look at masking the prompt out of the loss when fine-tuning small models for structured tasks like scoring or extraction.

fine-tuningllmunslothqlora