Parallel block drafters like DFlash guess every position in a block at once, which is fast but blind: no token sees its neighbours, so acceptance decays down the block, and verifying long blocks wastes batch capacity under load. DSpark (Cheng, Yu, Shao et al., DeepSeek-AI and Peking University) adds one fix per crack. Fix one, semi-autoregressive drafting: keep the parallel backbone but bolt on a tiny sequential head (a Markov or RNN transition bias) so each token conditions on the one before it, restoring coherence with almost no added latency, for +16 to 18% accepted length over DFlash and +27 to 31% over EAGLE-3. Fix two, confidence-scheduled verification: a confidence head predicts per-position survival odds, and a hardware-aware scheduler picks the verify length that maximizes system throughput tau times SPS(B) for the current load, verifying deep when idle and short when slammed. It stays exactly lossless. In production on DeepSeek-V4 it is 60 to 85% faster per user at matched throughput, holding up where the baseline hits a latency cliff. This is the fourth post in a series: scaled attention, speculative decoding, DFlash, DSpark.