papersTODAY 04:00 UTC
arXiv paper proposes training paradigm for fast long video generation
A new arXiv preprint addresses the difficulty of generating coherent minute-long videos, noting that while short clips are plentiful and high quality, long-form training data is scarce and confined to a few domains. The authors propose a training approach that combines mode-seeking and mean-seeking objectives to speed up long video generation. The work is positioned as a way to overcome the data bottleneck that limits scaling from seconds to minutes.