Akshath Tiwari

Speculative decoding still has a slow part: the draft model writes its guesses one token at a time, so drafting cost grows with the number of tokens guessed, and the best methods keep the drafter tiny to stay cheap, capping speedups around 2 to 3x. DFlash (Chen, Liang and Liu) unchains the two levers of the speedup equation L = (T_draft + T_verify) / tau. Move one: a block-diffusion model fills a whole block of masked tokens in a single parallel pass, so drafting cost stops growing with block size, which frees room for a deeper five-layer drafter with no latency penalty. Move two: it injects the target model’s hidden features into the Key and Value of every draft layer, so the guidance never fades with depth and acceptance keeps rising, unlike input-only conditioning. The result is over 6x lossless speedup on Qwen3-8B, about 2.5x beyond EAGLE-3, accepting roughly 6.5 tokens per cycle. This interactive essay covers the sequential-drafting cost, the two levers, parallel block drafting, KV-injected conditioning, benchmark results, batching behavior, and the reframing of diffusion models as disposable parallel scouts.