papersTODAY 04:00 UTC
Temporal Self-Distillation Speeds Up Discrete Diffusion Language Models
A new arXiv paper proposes Temporal Self-Distillation, a training method aimed at discrete diffusion language models that generate several tokens at once. Such models lose quality when too many tokens are decoded in parallel, and the technique is presented as a simple way to reduce that degradation. The approach targets faster inference without the accuracy drop that usually accompanies aggressive parallel decoding.