Continuous Diffusion Language Models Make a Comeback
Summary
Sander Dieleman traces the evolution of continuous diffusion language models from early embedding-based experiments to their 2026 revival. He highlights flow maps, adaptive schedules, self-conditioning, and hybrid approaches as efforts to improve scalability and sampling efficiency. The article argues that continuous models may be easier to distill into few-step or one-step generators while retaining token correlations, creating opportunities for faster inference and sampling-time steering. Discrete diffusion and autoregressive models remain strong alternatives, and the author expects all three paradigms to coexist. He also emphasizes that inconsistent evaluation protocols and surrogate generative-perplexity metrics make current comparisons difficult to interpret.