papersTODAY 04:00 UTC
AdaFlash: Adaptive Speculative Decoding with On-Policy Distilled Diffusion Drafters
A new arXiv paper proposes AdaFlash, a speculative decoding method that uses diffusion-based draft models distilled on-policy to speed up large language model inference. The approach adapts the drafting process rather than relying on a fixed draft model, aiming to improve acceptance rates during verification by the target model. It builds on prior work in this line, including DFlash, and appears as a revised submission.